altic-dev/nemotron-3-diarization-coreml
Nemotron 3 Diarization — Core ML
Core ML export of the final `nvidia/Nemotron-3-Diarization` checkpoint (revision 98fcee8e866bb534eee85d1eb3817170700ae717), for on-device speaker diarization on macOS 14+ / iOS 17+.
The license terms of the source checkpoint apply to this derivative.
Files
Precision: fp16 storage, fp32 compute
Weights are stored as fp16 and widened to fp32 at load time (constexpr_cast), so every op computes in fp32. The package stays at fp16 size (~199 MB).
- The checkpoint is bf16, which fp16 represents almost exactly (0.005% of values differ, by less than 3e-8).
- A plain fp16-compute export drifts on long audio: small per-chunk rounding feeds back through the speaker cache and compounds. On a 36-minute meeting it matched the reference on only 98.4–98.9% of frames and added a speaker. This happened on GPU, CPU and Neural Engine alike.
Validation
Frame-level agreement with NVIDIA's NeMo reference, same streaming loop and post-processing, on real meeting recordings:
Speed: ~21 ms per 27.2 s chunk on Apple silicon (about 2 s for a 36-minute track).
Interface
One call processes one chunk; the caller owns the speaker cache and FIFO update between calls (NeMo streaming_update).
Required runtime settings
These must match training, or the model splits one voice across several speaker slots:
spkcache_sil_frames_per_spk = 1- Fill silence slots in the speaker cache with
learnable_sil_emb.f32, not a running mean of silent frames. - Offline profile:
chunk_len 340,chunk_right_context 40,fifo_len 40,spkcache_len 264,spkcache_update_period 300.
