Team Ai
Modelpublic

altic-dev/nemotron-3-diarization-coreml

sourceHugging Faceotherupdated 13d agoView on Hugging Face
0likes908downloads
Model Card

Nemotron 3 Diarization — Core ML

Core ML export of the final `nvidia/Nemotron-3-Diarization` checkpoint (revision 98fcee8e866bb534eee85d1eb3817170700ae717), for on-device speaker diarization on macOS 14+ / iOS 17+.

The license terms of the source checkpoint apply to this derivative.

Files

FilePurpose
nemotron_3_diarization.mlpackageStreaming Sortformer step: pre-encoder, Fast-Conformer, transformer, speaker heads
learnable_sil_emb.f32512 raw little-endian float32 values: the trained silence embedding for the speaker cache
config.jsonStreaming geometry, cache and post-processing settings, I/O shapes, SHA-256 of every file
conversion/Scripts that produced the package

Precision: fp16 storage, fp32 compute

Weights are stored as fp16 and widened to fp32 at load time (constexpr_cast), so every op computes in fp32. The package stays at fp16 size (~199 MB).

  • —The checkpoint is bf16, which fp16 represents almost exactly (0.005% of values differ, by less than 3e-8).
  • —A plain fp16-compute export drifts on long audio: small per-chunk rounding feeds back through the speaker cache and compounds. On a 36-minute meeting it matched the reference on only 98.4–98.9% of frames and added a speaker. This happened on GPU, CPU and Neural Engine alike.

Validation

Frame-level agreement with NVIDIA's NeMo reference, same streaming loop and post-processing, on real meeting recordings:

AudioNeMo speakersCore ML speakersFrame agreement
70 s, one microphone22100.00%
36 min, microphone track4499.94%
36 min, application-audio track55100.00%

Speed: ~21 ms per 27.2 s chunk on Apple silicon (about 2 s for a 36-minute track).

Interface

One call processes one chunk; the caller owns the speaker cache and FIFO update between calls (NeMo streaming_update).

InputShapeType
chunk1 × 3040 × 128float32 log-mel, 10 ms hop
chunk_lengths1int32
spkcache1 × 264 × 512float32
spkcache_lengths1int32
fifo1 × 40 × 512float32
fifo_lengths1int32
OutputShape
spkcache_fifo_chunk_preds1 × 684 × 8
chunk_pre_encode_embs1 × 380 × 512
chunk_pre_encode_lengths1

Required runtime settings

These must match training, or the model splits one voice across several speaker slots:

  • —spkcache_sil_frames_per_spk = 1
  • —Fill silence slots in the speaker cache with learnable_sil_emb.f32, not a running mean of silent frames.
  • —Offline profile: chunk_len 340, chunk_right_context 40, fifo_len 40, spkcache_len 264, spkcache_update_period 300.