aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test
Nemotron 3.5 Core ML encoder export experiment
Experimental 320 ms Nemotron bundle for testing the ANE specialization failure on A16 described in speech-swift issue 503. Only the encoder conversion target changes from iOS 18 to iOS 17. The export preserves all 296 published encoder palettes and copies the decoder, joint network, and runtime metadata from the published bundle without changes.
This bundle does not lower the Swift SDK's iOS 18 requirement. A successful macOS load does not establish compatibility with the affected iPhone.
Files
Validation
M5 Pro, macOS 26.6.2 (25G83):
- All 296 encoder palette lookup tables and index arrays match the published values byte for byte. The decoder, joint, and runtime metadata are unchanged.
- All six encoder outputs match the published CPU encoder exactly across 67 calls with real speech, silence, partial chunks, and full streaming caches.
- Four SDK tests passed in each of three separate processes: published CPU+ANE, candidate CPU+ANE, and candidate CPU-only. Batch, streaming, and word-boosted transcripts matched the published result. The boosting engagement test also passed in all three configurations.
- The supplied probe compiles with Swift 6 on macOS and typechecks for arm64 iOS 18. The export-helper suite passed all 14 unit tests.
The SDK tests used speech-swift revision 7fc8f6c2b7847cad17641cf294b2854e20d936a8. Validation covers one English fixture; it does not establish multilingual accuracy or A16 compatibility.
Local load and encoder timing
The standalone Core ML probe ran one bundle and compute configuration per process. Prediction measurements use synthetic input, not full ASR.
Core ML system-cache state was uncontrolled, so first-load times are observations rather than a speedup claim. CPU+ANE compute plans for both encoders prefer the ANE for 1,592 operations and CPU for 86. Planned placement does not prove execution on the ANE. Preserve device compiler logs alongside the reports.
See testing/mac-validation.json for the measurements. Repeated loads and real-speech tests on the affected iPhone remain the acceptance check.
Usage
Download the snapshot into a separate directory and follow TESTING.md. On macOS, the Python Core ML loader can check the candidate encoder directly:
import coremltools as ct
from huggingface_hub import snapshot_download
bundle = snapshot_download(
"aufklarer/Nemotron-3.5-ASR-Streaming-0.6B-CoreML-INT8-iOS17-Encoder-Test",
local_dir="nemotron-ios17-encoder-test",
)
encoder = ct.models.CompiledMLModel(
f"{bundle}/encoder.mlmodelc",
compute_units=ct.ComputeUnit.CPU_AND_NE,
)To obtain a JSON load report on macOS:
swiftc -O -parse-as-library -D LOAD_PROBE_CLI nemotron-ios17-encoder-test/testing/LoadProbe.swift -o /tmp/nemotron-load-probe
/tmp/nemotron-load-probe nemotron-ios17-encoder-test ane candidate-ane.jsonWith speech-swift v0.0.27 or later, use the unchanged local-bundle API:
import CoreML
import NemotronStreamingASR
let model = try await NemotronStreamingASRModel.fromLocal(
bundleDir: candidateBundleURL,
computeUnits: .cpuAndNeuralEngine)Preserve the stock bundle for the baseline comparison. The model remains an experiment until the affected iPhone passes repeated-load and real-speech batch/streaming checks.
Source and license
Upstream: NVIDIA Nemotron 3.5 ASR Streaming 0.6B, revision f3d333391852ba876df169dcc9ba902d25b6ab0b.
Baseline: published Core ML INT8 bundle. The model uses the upstream OpenMDW 1.1 license.
