CodonProject/Whisper-Tiny-Encoder
0
1---2language:3- en4- multilingual5license: apache-2.06library_name: codon7pipeline_tag: feature-extraction8tags:9- audio10- speech11- whisper12- encoder13- feature-extraction14- safetensors15---16 17# Whisper-Tiny Audio Encoder (codon format)18 19Encoder-only weights of [openai/whisper-tiny](https://huggingface.co/openai/whisper-tiny),20re-keyed into the naming convention used by the21[codon](https://github.com/CodonProject/codon-model) library so they can be loaded into22`codon.block.model.WhisperTinyAudioEncoder` with a single call.23 24This is a **feature extractor**, not a speech-to-text model: the decoder and the25tokenizer are intentionally **not** included, so it cannot transcribe audio on its own.26 27## Files28 29| File | Size | Tensors | Description |30|---|---|---|---|31| `whisper_tiny_encoder.safetensors` | 15.7 MB | 67 | Encoder weights, **fp16**, codon key convention |32 33## Model details34 35| | |36|---|---|37| Architecture | Whisper encoder (pre-LN Transformer, bidirectional self-attention) |38| Parameters | 8,208,384 (8.21M) |39| `d_model` | 384 |40| Encoder layers | 4 |41| Attention heads | 6 (head dim 64) |42| FFN dim | 1536, GELU |43| Mel bins | 80 |44| Max positions | 1500 |45| Output | `[B, T/2, 384]` for input mel `[B, 80, T]` |46| Precision | fp16 (upcast to fp32 on load for fp32 execution) |47| License | Apache-2.0 |48 49## Requirements50 51```bash52pip install codon-model==0.0.7b853```54 55Source: [CodonProject/codon-model](https://github.com/CodonProject/codon-model)56 57## Usage58 59```python60import torch61from codon.block.model import WhisperTinyAudioEncoder62 63encoder = WhisperTinyAudioEncoder(pool_stride=1)64encoder.load('whisper_tiny_encoder.safetensors', strict=True)65encoder = encoder.float() # weights are fp16; upcast for fp32 execution66encoder.eval()67 68mel = torch.randn(1, 80, 3000) # 30 s of log-mel @ 100 Hz69with torch.no_grad():70 hidden, _ = encoder(mel)71print(hidden.shape) # torch.Size([1, 1500, 384])72```73 74Half-precision execution is also supported — just call `.half()` and feed a75half-precision mel instead (expect a larger numerical drift, ~6e-2, since the76arithmetic itself runs in fp16):77 78Loading straight from this repository:79 80```python81encoder = WhisperTinyAudioEncoder().from_remote()82```83 84### Downsampling for LLM consumption85 86The raw encoder emits one token per 20 ms (50 tokens/s), which is usually far too87dense for a language model. `pool_stride` appends a non-overlapping average pool88after the encoder (weights are unaffected):89 90```python91encoder = WhisperTinyAudioEncoder(pool_stride=8)92encoder = encoder.float()93with torch.no_grad():94 hidden, _ = encoder(torch.randn(1, 80, 3000))95print(hidden.shape) # torch.Size([1, 187, 384]) -> 6.25 tokens/s96```97 98## Tensor naming99 100Keys follow the codon convention (`proj_*` for projections). The full layout:101 102```103conv1.weight (384, 80, 3)104conv1.bias (384,)105conv2.weight (384, 384, 3)106conv2.bias (384,)107embed_positions.weight (1500, 384)108layers.{0..3}.attn_norm.weight (384,)109layers.{0..3}.attn_norm.bias (384,)110layers.{0..3}.attn.proj_q.weight (384, 384)111layers.{0..3}.attn.proj_q.bias (384,)112layers.{0..3}.attn.proj_k.weight (384, 384) # no bias113layers.{0..3}.attn.proj_v.weight (384, 384)114layers.{0..3}.attn.proj_v.bias (384,)115layers.{0..3}.attn.proj_o.weight (384, 384)116layers.{0..3}.attn.proj_o.bias (384,)117layers.{0..3}.fn_norm.weight (384,)118layers.{0..3}.fn_norm.bias (384,)119layers.{0..3}.mlp.proj_fc1.weight (1536, 384)120layers.{0..3}.mlp.proj_fc1.bias (1536,)121layers.{0..3}.mlp.proj_fc2.weight (384, 1536)122layers.{0..3}.mlp.proj_fc2.bias (384,)123norm.weight (384,)124norm.bias (384,)125```126 127Note that `proj_k` has **no bias** — this mirrors the original Whisper design128(its `k_proj` is the only projection without a bias in both the encoder and the129decoder). The positional table uses Whisper's own layout: `weight[:, :192]` holds130sine values and `weight[:, 192:]` holds cosine values for131`inv_freq = exp(-log(10000) * arange(192) / 192)`.132 133## Provenance and verification134 135Weights are copied from `openai/whisper-tiny` (`model.encoder.*`), renamed, and136saved as fp16. No values were modified: converting the fp32 export to fp16 and back137to fp32 is bit-exact element-wise, i.e. the source values already only carried fp16138precision (max absolute weight magnitude is ~16, well inside fp16 range).139 140The codon implementation was verified against the official `transformers`141implementation (`WhisperModel(...).encoder`) with both models loading the same142checkpoint (upcast to fp32) and consuming the same input:143 144| mel length | output shape | max abs difference |145|---|---|---|146| 3000 | `[2, 1500, 384]` | 5.9e-05 |147| 1000 (padded) | `[2, 500, 384]` | 3.9e-05 |148| 512 (padded) | `[2, 256, 384]` | 2.2e-05 |149 150Differences are at fp32 rounding level, not implementation differences.151 152## Limitations153 154- **Encoder only.** No decoder, no vocabulary, no tokenizer — cannot produce text.155- **Fixed 30 s input.** Whisper pads every input to 3000 mel frames. Short clips156 must be zero-padded to 3000 and the padded output frames discarded; the padded157 region does influence the valid region through convolution and LayerNorm, so it158 is not equivalent to running the encoder on a shorter input.159- **Not causal / not streaming.** The encoder is bidirectional and stateless, so it160 cannot be used for incremental streaming without recomputing a full window.161- **Precision.** Tensors are stored as fp16. The fp32 -> fp16 -> fp32 round trip is162 element-wise exact, so no information was lost relative to the source export, and163 running the model in fp32 after upcasting reproduces the fp32 reference exactly.164 Native fp16 execution (`model.half()`) is supported but introduces the usual165 half-precision drift (~6e-2 on this model).166 167## Citation168 169If you use these weights, please cite the original work:170 171```bibtex172@misc{radford2022whisper,173 doi = {10.48550/ARXIV.2212.04356},174 url = {https://arxiv.org/abs/2212.04356},175 author = {Radford, Alec and Kim, Jong Wook and Xu, Tao and Brockman, Greg and McLeavey, Christine and Sutskever, Ilya},176 title = {Robust Speech Recognition via Large-Scale Weak Supervision},177 publisher = {arXiv},178 year = {2022},179 copyright = {arXiv.org perpetual, non-exclusive license}180}181```182 183## Acknowledgements184 185Original model by OpenAI. This repository only re-keys the encoder weights for use186with [codon](https://github.com/CodonProject/codon-model).187 