Masterx/Audio8-ASR-Infinite-ONNX
Audio8-ASR-Infinite — ONNX (streaming, constant memory)
An ONNX Runtime export of [Edge0/Audio8-ASR-Infinite](https://huggingface.co/Edge0/Audio8-ASR-Infinite) (revision 7476824bc222e4ad509d286e8cae8b8d3f371129). All credit for the model, its training and its streaming design goes to Edge0. This repository only re-packages the original weights as split ONNX graphs with explicit state so the model can run natively streaming outside PyTorch. It was made for WinSTT, but the graphs are plain ONNX and the IO contract is documented below.
Licence: Apache-2.0, inherited from the original model. See LICENSE.
What the model is
- A causal Voxtral-Realtime-style audio tower with 32 layers, hidden size 1280, 32 heads × 64, and sliding-window attention over 750 frames. The log-mel and the causal conv front end are inside the graph.
- A projector into a Qwen2.5-3B decoder with 36 layers, 16 query heads and 2 KV heads, head_dim 128, and a tied LM head. Voxtral's AdaRMSNorm delay modulation is applied per layer.
- Semantic-VAD heads on the decoder's last hidden state: 4 horizons (0.5 / 1 / 2 / 3 s) × 8 classes. Class 0 at the 2 s horizon is the end-of-turn probability.
- About 4.09B parameters in total. Languages: Chinese and English, selected by a prompt token.
- Streaming clock. One text token per 1280 samples (80 ms), with a fixed 480 ms delay (6 tokens). Text appears word by word while audio is still arriving. Audio length is unbounded: the encoder keeps a 749-slot ring buffer and the decoder keeps a 375-slot rolling cache. When the decoder cache fills, the oldest 38 entries are dropped after a 16-token stable prefix. Memory therefore stays constant however long the input is.
Files
Download sizes per precision
Each size includes the shared tables and tokenizer (636 MB).
In every tier, only the linear weights are quantized; activations stay in fp32 (fp16 for the fp16 tier).
Graph IO contract
The rotary embedding is applied on the host. Keys are cached already rotated, using angles computed in f64 and reduced mod 2π. Because of that, no cache ever needs re-rotation, except for the decoder's rolling trim, where the surviving suffix is re-based by −38 positions.
*`audio_encoder_.onnx`**
*`decoder_.onnx`**
Decode loop. It follows the original simulated_streaming_greedy.
- Window 0 covers audio
[0, 25·1280 + 40). Pass it to the encoder with prefill 25. - The decoder prompt is
[BOS, <lang>, STREAMING_PAD × 23], added to the 25 audio embeddings. - Every later 80 ms window yields one audio embedding and exactly one decoder step. The step's input is the previously emitted token's embedding plus that audio embedding.
- Pick the next token by greedy argmax, with EOS masked out.
The audio is left-padded by 18 tokens. At end of stream, add (6 + 1 + 10) × 1280 samples of silence. Visible text is every token below id 151643, detokenized with byte-level BPE. scripts/audio8_infinite_ref.py (class Session) is a complete NumPy reference of this loop.
Parity and accuracy
Setup
- Reference: the original PyTorch weights in fp32. Decoding is exact greedy with EOS masked, following the same streaming schedule. The checkpoint is read directly by
scripts/audio8_infinite_ref.py, sotransformersis not involved. - Logit and token parity: measured on JFK (11 s, 148 decoder steps) and on JFK ×6 (66 s, 835 steps). The 66 s clip overflows the 375-slot decoder cache, so it goes through 13 rolling trims.
- Accuracy: measured with the WinSTT Rust runtime on two concatenated sets, each streamed as one continuous recording of about 5 minutes:
- 40 LibriSpeech test-clean utterances (308 s), scored by WER;
- 25 FLEURS cmnhanscn utterances (290 s), scored by CER.
- Scoring normalization: English is lower-cased with punctuation stripped. Chinese is NFKC-normalized with punctuation and whitespace stripped; digits and Chinese numerals are not unified, which penalizes all models equally.
Parity against PyTorch fp32 (JFK, 148 steps)
The peak logit magnitude is about 37.
Long audio (66 s, 835 steps, 13 rolling-cache trims):
- The
int4sequence of visible text tokens is identical to the PyTorch sequence (156/156 tokens), and the transcripts are identical. - The only differences are in the timing of the non-visible pad/word markers.
Accuracy (WinSTT Rust runtime, CPU, continuous streaming of the concatenated sets)
Notes on the comparison rows:
- Audio8-ASR-0.1B and ARK-ASR-0.6B are offline models. WinSTT runs them on VAD-split chunks of up to 24 s with full right context.
- Audio8-ASR-Infinite decodes the same audio strictly causally, with 480 ms of look-ahead.
- These are small subsets, so treat differences under about 1 point as noise.
Set the language token. Prompting with the wrong language makes the model translate. With the English token, Chinese speech came out as English. At int8 this also degenerated into a repetition loop. The model cannot auto-detect the language, so pass zh or en explicitly.
Speed
All timings below were taken on a Windows desktop, a 24-thread CPU, while other heavy jobs were running. Use them as relative numbers, not absolute ones.
- CPU, 11 s clip, same moment (MatMulNBits
int4): RTF 2.1 on CPU vs 10.8 with DirectML (RTX 3080 Ti). The transcripts were identical. - DirectML. The streaming caches are re-fed from the host on every call: about 390 MB of encoder ring per window and about 27 MB of decoder cache per token. That makes DirectML roughly 5× slower than CPU here, so WinSTT pins this model to CPU.
- CPU across precisions. RTF ranged from 1.0 to 4.2 for
int4/int8over the 5-minute sets, depending on load. The two tiers were within the load noise of each other, soint8is the recommended tier: it is more accurate and not measurably slower. - fp16 is slow on CPU: RTF 20.6 on the 290 s FLEURS set. It is included for GPU runtimes; WinSTT itself offers only
int8andint4. - Load and memory. Load takes about 15 s. Peak RSS is about 5.4 GB at
int4, and memory stays constant regardless of audio length.
Reproducing
python scripts/audio8_infinite_export.py --ckpt <snapshot of Edge0/Audio8-ASR-Infinite> --out bundle --precision int4 int8 fp16
A8I_CKPT=<snapshot> A8I_BUNDLE=bundle python scripts/audio8_infinite_parity.py torch --sets jfk
A8I_CKPT=<snapshot> A8I_BUNDLE=bundle python scripts/audio8_infinite_parity.py onnx --prec int4 --sets jfk ls_clean fleurs_zh --keep-logits
python scripts/audio8_infinite_parity.py reportThe scripts need onnx, onnxruntime (≥ 1.20 for 8-bit MatMulNBits), numpy, and, for the parity harness, tokenizers, soundfile and torch. The checkpoint is read straight from the safetensors files, so no transformers install is needed. They were built with onnx 1.23.1, onnxruntime 1.30.0 and torch 2.13.0.
Credits
- Model, training and streaming design: Edge0, Edge0/Audio8-ASR-Infinite, Apache-2.0.
- The architecture builds on Mistral's Voxtral-Realtime audio tower and Alibaba's Qwen2.5-3B.
- ONNX export and streaming runtime: WinSTT.
