Team Ai
Modelpublic

Masterx/Audio8-ASR-Infinite-ONNX

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes82downloads
Model Card

Audio8-ASR-Infinite — ONNX (streaming, constant memory)

An ONNX Runtime export of [Edge0/Audio8-ASR-Infinite](https://huggingface.co/Edge0/Audio8-ASR-Infinite) (revision 7476824bc222e4ad509d286e8cae8b8d3f371129). All credit for the model, its training and its streaming design goes to Edge0. This repository only re-packages the original weights as split ONNX graphs with explicit state so the model can run natively streaming outside PyTorch. It was made for WinSTT, but the graphs are plain ONNX and the IO contract is documented below.

Licence: Apache-2.0, inherited from the original model. See LICENSE.

What the model is

  • —A causal Voxtral-Realtime-style audio tower with 32 layers, hidden size 1280, 32 heads × 64, and sliding-window attention over 750 frames. The log-mel and the causal conv front end are inside the graph.
  • —A projector into a Qwen2.5-3B decoder with 36 layers, 16 query heads and 2 KV heads, head_dim 128, and a tied LM head. Voxtral's AdaRMSNorm delay modulation is applied per layer.
  • —Semantic-VAD heads on the decoder's last hidden state: 4 horizons (0.5 / 1 / 2 / 3 s) × 8 classes. Class 0 at the 2 s horizon is the end-of-turn probability.
  • —About 4.09B parameters in total. Languages: Chinese and English, selected by a prompt token.
  • —Streaming clock. One text token per 1280 samples (80 ms), with a fixed 480 ms delay (6 tokens). Text appears word by word while audio is still arriving. Audio length is unbounded: the encoder keeps a 749-slot ring buffer and the decoder keeps a 375-slot rolling cache. When the decoder cache fills, the oldest 38 entries are dropped after a 16-token stable prefix. Memory therefore stays constant however long the input is.

Files

FilePurpose
audio_encoder_{int4,int8,fp16}.onnx (+ .onnx.data)mel + causal conv + 32-layer tower + projector, one call per batch of new 80 ms windows
decoder_{int4,int8,fp16}.onnx (+ .onnx.data)Qwen2.5-3B step with KV cache in/out, last-position logits + VAD logits
embed_tokens.bf16raw bf16 [151936, 2048] token-embedding table (looked up on the host)
ada_scale.f32precomputed (1 + ada_rms_norm(t_cond)) table [8, 36, 2048], one row per (frame_len, delay) combo listed in runtime.json
runtime.jsonformat winstt-audio8-infinite-v1: streaming geometry, cache sizes, special tokens, VAD horizons, table layouts
tokenizer.json, tokenizer_config.json, config.jsonfrom the original repo
scripts/the export, graph builder, PyTorch reference and parity harness used to make this repo

Download sizes per precision

Each size includes the shared tables and tokenizer (636 MB).

PrecisionEncoderDecoderTotal
int4 (MatMulNBits 4-bit, block 32, asymmetric, accuracy_level 4)0.66 GB1.98 GB3.27 GB
int8 (MatMulNBits 8-bit, block 32, accuracy_level 4)1.17 GB3.57 GB5.37 GB
fp16 (whole graph fp16; RMSNorm, softmax and mel in fp32; fp16 KV caches)1.99 GB6.17 GB8.80 GB

In every tier, only the linear weights are quantized; activations stay in fp32 (fp16 for the fp16 tier).

Graph IO contract

The rotary embedding is applied on the host. Keys are cached already rotated, using angles computed in f64 and reduced mod 2π. Because of that, no cache ever needs re-rotation, except for the decoder's rolling trim, where the surviving suffix is re-based by −38 positions.

*`audio_encoder_.onnx`**

NameType / shapeMeaning
in audiof32 [B, S]B raw 16 kHz windows of equal length (one per new token, with 52.5 ms look-back / 2.5 ms look-ahead)
in frame_leni64 [1]encoder frames per text token (4 = 80 ms)
in cos, sinf32 [T, 64]RoPE factors of the T = B·4 new frames
in attn_biasf32 [T, 749 + T]additive mask (0 / −1e9): ring slots, then the new frames
in past_key_i, past_value_i[1, 32, 749, 64] ×32 layersencoder ring (rotated keys)
out audio_embedsf32 [1, T/4, 2048]one audio embedding per text token
out key_i, value_i[1, 32, T, 64]new keys and values to write into the ring

*`decoder_.onnx`**

NameType / shapeMeaning
in inputs_embedsf32 [1, n, 2048]embed_tokens[prev_token] + audio_embeds[k]
in ada_scalef32 [36, 2048]the row of ada_scale.f32 for the session's (frame_len, delay)
in cos, sinf32 [n, 128]RoPE factors of the new positions
in attn_biasf32 [n, 375 + n]additive mask: cache slots, then causal over the new positions
in past_key_i, past_value_i[1, 2, 375, 128] ×36 layersrolling decoder cache (fp16 in the fp16 tier)
out logitsf32 [1, 151936]last position
out vad_logitsf32 [1, 4, 8]semantic-VAD heads, last position
out key_i, value_i[1, 2, n, 128]new cache entries

Decode loop. It follows the original simulated_streaming_greedy.

  1. 1.Window 0 covers audio [0, 25·1280 + 40). Pass it to the encoder with prefill 25.
  2. 2.The decoder prompt is [BOS, <lang>, STREAMING_PAD × 23], added to the 25 audio embeddings.
  3. 3.Every later 80 ms window yields one audio embedding and exactly one decoder step. The step's input is the previously emitted token's embedding plus that audio embedding.
  4. 4.Pick the next token by greedy argmax, with EOS masked out.

The audio is left-padded by 18 tokens. At end of stream, add (6 + 1 + 10) × 1280 samples of silence. Visible text is every token below id 151643, detokenized with byte-level BPE. scripts/audio8_infinite_ref.py (class Session) is a complete NumPy reference of this loop.

Parity and accuracy

Setup

  • —Reference: the original PyTorch weights in fp32. Decoding is exact greedy with EOS masked, following the same streaming schedule. The checkpoint is read directly by scripts/audio8_infinite_ref.py, so transformers is not involved.
  • —Logit and token parity: measured on JFK (11 s, 148 decoder steps) and on JFK ×6 (66 s, 835 steps). The 66 s clip overflows the 375-slot decoder cache, so it goes through 13 rolling trims.
  • —Accuracy: measured with the WinSTT Rust runtime on two concatenated sets, each streamed as one continuous recording of about 5 minutes:
  • —40 LibriSpeech test-clean utterances (308 s), scored by WER;
  • —25 FLEURS cmnhanscn utterances (290 s), scored by CER.
  • —Scoring normalization: English is lower-cased with punctuation stripped. Chinese is NFKC-normalized with punctuation and whitespace stripped; digits and Chinese numerals are not unified, which penalizes all models equally.

Parity against PyTorch fp32 (JFK, 148 steps)

PrecisionMax \Δlogit\Mean per-step max \Δlogit\Greedy tokens identicalTranscript identical
fp32 graph (reduced-size test config)2.6e-6n/ayesyes
fp160.0470.014yes, 148/148yes
int82.830.87yes, 148/148yes
int44.90 (before the first divergence)2.81no: 97.97% agree, first split at step 46 on a [STREAMING_PAD] vs [STREAMING_WORD] boundary markeryes
legacy int8dq (not shipped)n/an/ano: split at step 7no ("For my fellow Americans, ask not. What…")

The peak logit magnitude is about 37.

Long audio (66 s, 835 steps, 13 rolling-cache trims):

  • —The int4 sequence of visible text tokens is identical to the PyTorch sequence (156/156 tokens), and the transcripts are identical.
  • —The only differences are in the timing of the non-visible pad/word markers.

Accuracy (WinSTT Rust runtime, CPU, continuous streaming of the concatenated sets)

ModelPrecisionLibriSpeech test-clean (40 utts) WERFLEURS zh (25 utts) CER
Audio8-ASR-Infiniteint84.27%11.25%
Audio8-ASR-Infiniteint44.66%11.85%
Audio8-ASR-Infinitefp16not run (too slow on CPU)11.37%
Audio8-ASR-0.1B (for comparison, chunked offline)fp323.75%12.09%
ARK-ASR-0.6B (for comparison, chunked offline)int82.72%11.00%

Notes on the comparison rows:

  • —Audio8-ASR-0.1B and ARK-ASR-0.6B are offline models. WinSTT runs them on VAD-split chunks of up to 24 s with full right context.
  • —Audio8-ASR-Infinite decodes the same audio strictly causally, with 480 ms of look-ahead.
  • —These are small subsets, so treat differences under about 1 point as noise.

Set the language token. Prompting with the wrong language makes the model translate. With the English token, Chinese speech came out as English. At int8 this also degenerated into a repetition loop. The model cannot auto-detect the language, so pass zh or en explicitly.

Speed

All timings below were taken on a Windows desktop, a 24-thread CPU, while other heavy jobs were running. Use them as relative numbers, not absolute ones.

  • —CPU, 11 s clip, same moment (MatMulNBits int4): RTF 2.1 on CPU vs 10.8 with DirectML (RTX 3080 Ti). The transcripts were identical.
  • —DirectML. The streaming caches are re-fed from the host on every call: about 390 MB of encoder ring per window and about 27 MB of decoder cache per token. That makes DirectML roughly 5× slower than CPU here, so WinSTT pins this model to CPU.
  • —CPU across precisions. RTF ranged from 1.0 to 4.2 for int4 / int8 over the 5-minute sets, depending on load. The two tiers were within the load noise of each other, so int8 is the recommended tier: it is more accurate and not measurably slower.
  • —fp16 is slow on CPU: RTF 20.6 on the 290 s FLEURS set. It is included for GPU runtimes; WinSTT itself offers only int8 and int4.
  • —Load and memory. Load takes about 15 s. Peak RSS is about 5.4 GB at int4, and memory stays constant regardless of audio length.

Reproducing

bash
python scripts/audio8_infinite_export.py --ckpt <snapshot of Edge0/Audio8-ASR-Infinite> --out bundle --precision int4 int8 fp16
A8I_CKPT=<snapshot> A8I_BUNDLE=bundle python scripts/audio8_infinite_parity.py torch --sets jfk
A8I_CKPT=<snapshot> A8I_BUNDLE=bundle python scripts/audio8_infinite_parity.py onnx --prec int4 --sets jfk ls_clean fleurs_zh --keep-logits
python scripts/audio8_infinite_parity.py report

The scripts need onnx, onnxruntime (≥ 1.20 for 8-bit MatMulNBits), numpy, and, for the parity harness, tokenizers, soundfile and torch. The checkpoint is read straight from the safetensors files, so no transformers install is needed. They were built with onnx 1.23.1, onnxruntime 1.30.0 and torch 2.13.0.

Credits

  • —Model, training and streaming design: Edge0, Edge0/Audio8-ASR-Infinite, Apache-2.0.
  • —The architecture builds on Mistral's Voxtral-Realtime audio tower and Alibaba's Qwen2.5-3B.
  • —ONNX export and streaming runtime: WinSTT.