spybyscript/nemotron-3.5-asr-streaming-0.6b-litert
Nemotron 3.5 ASR Streaming 0.6B — LiteRT conversion
Community conversion by spybyscript, derived from NVIDIA's official Nemotron 3.5 ASR model. This is not an official NVIDIA or Google release. No additional training or fine-tuning was performed.
These are component graphs for a stateful LiteRT integration, not a single drop-in transcription model. Each version contains 31 .tflite files plus a tensor/state manifest, vocabulary, source provenance and parity evidence. An application must implement audio chunking, cache management and greedy RNN-T decoding. The integration source repository is private; no APK or private implementation source is distributed here. The public integration contract describes how to use these files independently.
Punctuation correction — 2026-09-25
The current bundles use `nemotron35-hf560-v2` and restore missing sentence-ending punctuation by masking the final padded chunk through subsampling and encoder attention. All original constant payloads and quantization parameters are preserved. Full-chunk outputs and caches were bit-identical in the preservation tests. Added graph size is only about 39–57 KB per variant.
A host update is required: subsampling graphs add Int32 args_4 (valid current mel frames), every encoder block adds Int32 args_5 (valid current encoder frames), and the decoder must consume only valid final frames. See INTEGRATION.md. Old hosts can pin revision 153a7dbd809a255e5999130397bcd60a66f81a54.
Across 80 clips from 40 speakers, all variants went from 0 to 73 transcripts with terminal punctuation, with zero changes to recognized words versus the same-precision previous bundle. Corrected FP32 matched the original safetensors model on all 80 transcripts; FP16 matched 79 and INT8 matched 68. The original eight diagnostic clips, including two questions, regained punctuation in all variants. No punctuation is inserted by post-processing.
See PUNCTUATION_FIX.md, validation-summary.json, and each variant's validation reports for scope and retained quantization differences. The multi-speaker test used Python LiteRT 2.1.2 CPU; 24 diagnostic cases per variant also passed on LiteRT 2.2.0. The updated graphs have not been rerun on Android or GPU/NPU hardware.
A standalone CPU reference runner is now included. It loads only the chosen bundle and implements the new finalization contract. The original private integration source is still not included.
Source and license
- Base model: nvidia/nemotron-3.5-asr-streaming-0.6b.
- Exact source revision: `1c8deaecc64b91f034d73e08dd8b64625eb3395d`.
- Original weights SHA-256:
9eebdd6590289cb3030f310858f3df93256600a800a3e8200c5993d5f967e174. - Model materials: OpenMDW-1.1 license; official license page. Retain the license and applicable origin notices when redistributing.
- See NOTICE.md and each folder's
source-manifest.json. EachMODEL_CARD.mdis an unchanged copy of the upstream card, retained for attribution; its benchmarks and multilingual claims are not results of this conversion.
The upstream checkpoint is multilingual. This release's tested integration supports en-US, prompt ID 0, and the 560 ms chunk profile only. Other languages, automatic language detection and other chunk sizes have not been validated here. Do not infer their support from the upstream card or vocabulary.
Versions
INT8 is partial quantization, not a fully integer model: 192 fully-connected weight buffers across 24 encoder blocks are INT8; 72 convolution/depthwise weight buffers remain FP32. Non-encoder-block graphs remain FP32 and byte-identical to the FP32 baseline. Activations/interfaces and streaming caches are floating point. No representative-dataset static activation calibration was used. This is not a precompiled Qualcomm NPU model.
The FP16 variant includes a bounded attention-mask rewrite (-infinity to -10000) used in the GPU experiments. FP16 storage does not guarantee that every CPU operation executes with FP16 arithmetic. FP32 and INT8 retain their original negative-infinity mask values; v2 adds current-frame validity masking to all variants.
Historical v1 Android accuracy and performance
These measurements apply to the previous v1 revision, not a new benchmark of v2. Samsung S23 Ultra, Snapdragon 8 Gen 2, Android 16, LiteRT 2.2.0, four CPU threads. Small development measurements, not a standard full benchmark. Human audio was 23 clean English audiobook clips from 10 speakers, totaling 569 reference words. Three synthetic clips were evaluated separately. Corpus variant order rotated per clip, with a new app process for each run.
RTF is measured compute seconds divided by audio seconds; lower is faster and below 1 means faster than realtime in that test. It is not first-text latency. Model loading is excluded. INT8 added five word errors; equal FP32/FP16 aggregate error counts do not mean every transcript was identical. All successful runs had finite outputs and no RNN-T symbol-limit failures.
The sustained test used 574.935 seconds of concatenated/repeated audio without resetting model state. It is not additional independent accuracy data. It ran INT8 → FP16 → FP32 without cooling, with thermal statuses respectively 1–2, 2–3 and 3. These differing conditions limit direct speed comparisons. Resident PSS subtracts same-snapshot SwapPss; the later median excludes startup and teardown. It is neither model-file size nor an exact peak. The long replay also held an audio buffer in memory.
The corpus used intrusive full process memory dumps. An initial sustained FP32 attempt under that method was killed for LOW_MEMORY; all three variants completed after switching to local-only memory sampling approximately every 5–6 seconds. This does not establish that FP32 cannot run long sessions or isolate the cause of the earlier failure. These are accelerated replays, not hours-long microphone-paced battery tests. TTS/LLM coexistence, energy consumption and broad noisy/accent/domain accuracy remain untested.
INT8 acceptance and historical parity evidence
INT8 changes source outputs and failed the original exact-token/numerical smoke-parity gate. The old parity-*.json files and previous-manifest.json remain unchanged as historical evidence for v1; current repair checks are in punctuation-validation.json and preservation-validation.json. Its parity-*.json reports remain unchanged, including passed: false; its manifest retains evaluation_only: true. The previous development harness explicitly accepted only the following v1 manifest after its on-device comparison. That acceptance does not automatically apply to the new v2 manifest:
0b5b594dcd499de6ffc7b2511ade5d493bb3b93a99e311192c64b1fc1ae58405
This acceptance records a chosen accuracy/memory tradeoff, not parity equivalence. A generic loader that rejects evaluation bundles will still reject it. Do not edit the manifest or mark failed reports as passed to make it load. An independent integration should implement its own explicit acceptance policy while preserving file hashes and evidence.
FP32 manifest SHA-256: 9be792bf2a6ee961020b86f29cfe156591dd884a9aedb96d239282f861692ca9.
FP16 manifest SHA-256: 0ea7b854d815682a48ec7eb3632b3a4ec8d2af214f0d51d86e5b986bdb98def1.
Current v2 INT8 manifest SHA-256: 996b26e5069bf0b921c2ade14ce10c7bdccee075300a184794554706e27075c8. It retains evaluation_only: true; integrations must explicitly evaluate/accept this identity and implement v2 inputs.
Download and integrate
Download one complete variant, plus the root license, notices and documentation. With huggingface_hub installed:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="spybyscript/nemotron-3.5-asr-streaming-0.6b-litert",
# Pin revision to an immutable commit from the repository history for deployment.
allow_patterns=["int8/*", "README.md", "INTEGRATION.md", "LICENSE",
"NOTICE.md", "publication.json", "SHA256SUMS", "PUNCTUATION_FIX.md",
"validation-summary.json", "reference_inference.py"],
local_dir="nemotron-litert",
)Verify the manifest identity above and every inventoried file's SHA-256 and byte count before loading. SHA256SUMS covers all published payloads except itself; publication.json identifies the original conversion source revision and bundle manifests. Keep each variant in its own directory; do not mix graphs or vocabularies from different bundles/models.
Use the integration contract with the actual input/output names, shapes and dtypes in each manifest. LiteRT Android's standard Kotlin CompiledModel API was used for inference. CPU4 is the tested INT8 configuration. GPU experiments required explicit FP32 arithmetic and mixed CPU/GPU execution on FP16 graphs, used more memory, and established no general sustained advantage. NPU acceleration has not been implemented or verified. These files do not load directly with transformers.pipeline or nemo.from_pretrained, and this repository does not provide hosted inference.
Limitations and intended use
For local speech recognition experiments and application integration. Errors can include names, word substitutions, omissions and punctuation. The small clean-speech evaluation does not establish performance in meetings, noise, varied accents or specialist terminology. Upstream training, bias, intended-use and ethical information is available in the pinned NVIDIA model card. No claim of government-contract suitability or certification is made. The model download is public; the private development repository is not required to download or inspect the artifacts.
Bundle manifest filename and completeness
Each precision folder uses nemotron-manifest.json. An earlier v1 update changed only that filename; this v2 update changes the graph input contract and manifest identities as described above. Renaming a v2 manifest does not make it compatible with a v1 host. Pin 153a7dbd809a255e5999130397bcd60a66f81a54 for v1.
Each of fp32/, fp16/, and int8/ is a self-contained model-data bundle. Download the chosen folder's manifest and every file in its inventory. The bundle includes all 31 LiteRT graphs (including the audio frontend), vocabulary with special token IDs, and validation/provenance files. No extra tokenizer, NVIDIA checkpoint, external weight file, or model download from another repository is required for inference. Source URLs in provenance files provide attribution and reproducibility; the runtime does not fetch them.
Applications still need the LiteRT runtime and the streaming/greedy RNN-T orchestration described in INTEGRATION.md. These model files are not a standalone speech SDK. After installing the bundle and runtime, transcription can run offline.
