Team Ai
Modelpublic

aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test

sourceHugging Faceopenmdw-1.1updated 4h agoView on Hugging Face
0likes
Model Card

Nemotron 3.5 Core ML encoder export experiment

Experimental 320 ms Nemotron bundle for testing the ANE specialization failure on A16 described in speech-swift issue 503. Only the encoder conversion target changes from iOS 18 to iOS 17. The export preserves all 296 published encoder palettes and copies the decoder, joint network, and runtime metadata from the published bundle without changes.

PropertyValue
ArchitectureCache-aware FastConformer encoder and RNN-T decoder
Parameters600 million
FormatCompiled Core ML .mlmodelc
Encoder weightsPublished 8-bit palette indices and FP16 lookup tables
Decoder and joint weightsFP16, unchanged
Runtime file size613 MiB
Audio16 kHz mono
Streaming chunk320 ms
Encoder minimum deployment targetiOS 17
Decoder and joint minimum deployment targetiOS 18, unchanged
Baseline revision447095fe87b480b5e6a15367f135303d479de8ac
StatusExperimental; affected-iPhone validation pending

This bundle does not lower the Swift SDK's iOS 18 requirement. A successful macOS load does not establish compatibility with the affected iPhone.

Files

FileSizePurpose
encoder.mlmodelc/565.4 MiBEncoder compiled for the iOS 17 operation set
decoder.mlmodelc/28.5 MiBUnchanged published RNN-T decoder
joint.mlmodelc/18.0 MiBUnchanged published joint network
config.json589 BUnchanged streaming geometry
vocab.json230.6 KiBUnchanged vocabulary
languages.json2.0 KiBUnchanged language map
tokenizer.model397.0 KiBUnchanged SentencePiece tokenizer
experiment.jsonSmall JSON; exact inventory includedSource revisions, export versions, and artifact hashes
TESTING.md6.5 KiBPublished versus candidate A/B instructions
testing/LoadProbe.swift7.8 KiBStandalone iOS/macOS Core ML load probe
testing/test_audio.wav625.0 KiBReal-speech smoke-test fixture, 20 s, mono 16 kHz
testing/mac-validation.json30.9 KiBLocal load, encoder-output, and SDK test results
testing/fixture-provenance.json<1 KiBOriginal fixture and resampling hashes

Validation

M5 Pro, macOS 26.6.2 (25G83):

  • —All 296 encoder palette lookup tables and index arrays match the published values byte for byte. The decoder, joint, and runtime metadata are unchanged.
  • —All six encoder outputs match the published CPU encoder exactly across 67 calls with real speech, silence, partial chunks, and full streaming caches.
  • —Four SDK tests passed in each of three separate processes: published CPU+ANE, candidate CPU+ANE, and candidate CPU-only. Batch, streaming, and word-boosted transcripts matched the published result. The boosting engagement test also passed in all three configurations.
  • —The supplied probe compiles with Swift 6 on macOS and typechecks for arm64 iOS 18. The export-helper suite passed all 14 unit tests.

The SDK tests used speech-swift revision 7fc8f6c2b7847cad17641cf294b2854e20d936a8. Validation covers one English fixture; it does not establish multilingual accuracy or A16 compatibility.

Local load and encoder timing

The standalone Core ML probe ran one bundle and compute configuration per process. Prediction measurements use synthetic input, not full ASR.

Bundle / compute unitsFirst observed encoder loadSame-process reloadMedian repeated encoder prediction
Published / CPU+ANE10.53 s76.5 ms8.09 ms
Candidate / CPU+ANE, first process6.63 s62.7 ms9.10 ms
Candidate / CPU+ANE, second process130.0 ms65.4 ms8.45 ms
Candidate / CPU-only2.77 s53.1 ms15.50 ms

Core ML system-cache state was uncontrolled, so first-load times are observations rather than a speedup claim. CPU+ANE compute plans for both encoders prefer the ANE for 1,592 operations and CPU for 86. Planned placement does not prove execution on the ANE. Preserve device compiler logs alongside the reports.

See testing/mac-validation.json for the measurements. Repeated loads and real-speech tests on the affected iPhone remain the acceptance check.

Usage

Download the snapshot into a separate directory and follow TESTING.md. On macOS, the Python Core ML loader can check the candidate encoder directly:

python
import coremltools as ct
from huggingface_hub import snapshot_download

bundle = snapshot_download(
    "aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test",
    local_dir="nemotron-ios17-encoder-test",
)
encoder = ct.models.CompiledMLModel(
    f"{bundle}/encoder.mlmodelc",
    compute_units=ct.ComputeUnit.CPU_AND_NE,
)

To obtain a JSON load report on macOS:

bash
swiftc -O -parse-as-library -D LOAD_PROBE_CLI nemotron-ios17-encoder-test/testing/LoadProbe.swift -o /tmp/nemotron-load-probe
/tmp/nemotron-load-probe nemotron-ios17-encoder-test ane candidate-ane.json

With speech-swift v0.0.27 or later, use the unchanged local-bundle API:

swift
import CoreML
import NemotronStreamingASR

let model = try await NemotronStreamingASRModel.fromLocal(
    bundleDir: candidateBundleURL,
    computeUnits: .cpuAndNeuralEngine)

Preserve the stock bundle for the baseline comparison. The model remains an experiment until the affected iPhone passes repeated-load and real-speech batch/streaming checks.

Source and license

Upstream: NVIDIA Nemotron 3.5 ASR Streaming 0.6B, revision f3d333391852ba876df169dcc9ba902d25b6ab0b.

Baseline: published Core ML INT8 bundle. The model uses the upstream OpenMDW 1.1 license.

Links