coder543/granite-speech-4.1-2b-nar-glade
Granite Speech 4.1 2B NAR for Glade
Converted model assets and runtime configuration for Glade.
A community conversion of IBM Granite Speech 4.1 2B NAR for Glade, using Apple's Core AI APIs and optimized for Neural Engine execution on Apple silicon. The model combines a CTC acoustic encoder with one bidirectional, audio-conditioned transcript-editing pass.
This is a W8A16 export: most convolution/projection weights are INT8 and activations are FP16. Conformer convolution output projections, normalization parameters, and the embedding table remain FP16. Graphs use ANE-oriented layouts, fused layers, and shared weights across a bounded set of static input shapes. Debug source metadata has been removed.
The repository contains 59 .aimodel components, embeddings.f16, tokenizer.json, model configuration in metadata.json. The components cover the encoder, acoustic projector, CTC head, editor, and final vocabulary head. These are portable source assets; Core AI specializes them for the target device on first use, which can take several minutes.
Requirements: Apple silicon with macOS 27 or a physical iOS 27 device supporting Core AI. These assets require application-side orchestration, audio preprocessing, CTC collapse, and tokenization; they are not a drop-in Transformers checkpoint. No inference runtime is included here. Use 16 kHz mono audio and bounded chunks (60 seconds or shorter) for longer recordings. The model is intended for final transcription, rather than native streaming.
Converted from upstream revision `a1e3416e25ce29ab3852778e54fa8b3bd59c4bf2`. Original model by IBM; this community conversion is not an official IBM or Apple release. Distributed under Apache-2.0, matching the upstream model. See the upstream model card for training details and limitations.
Measured performance
M3 MacBook Air (16 GB), macOS 27. The audio is JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the full 18-minute 15-second recording. The qualified host uses shared encoder/editor shapes and bounded chunks of at most 60 seconds, with CPU Silero choosing boundaries. Silero is not included in this repository.
Medians of three runs after preparation and warmup, including audio reading, chunk planning, the frontend, encoder, projector, editor and decoding. Preparation is excluded. Both measurements used macOS thermal state “fair” after compilation. The full recording retains all samples across 21 chunks. Benchmark-normalized WER is 52/2,219 (2.34%) against the supplied reference; this is one recording, not a general or multilingual accuracy evaluation. A separate full-recording trace recorded 1,844 ANE predictions and zero GPU intervals. Placement is not a measure of arithmetic utilization.
Device specialization and loading
For the same 59-component shared-shape source configuration on this Mac:
Preparation includes model loading and compilation requested by Core AI, before warmup or transcription. These observations do not isolate a clean first-install compile or the incremental cost of adding shapes. They preceded removal of authoring debug metadata, which preserves computation but changes artifact identity. Underlying ANE cache state is not fully observable, and cache loss or OS/device/application changes can require specialization again. These Mac measurements are not iPhone startup estimates.
