Team Ai
Modelpublic

aufklarer/ReDimNet2-B6-CoreML

sourceHugging Facemitupdated 5d agoView on Hugging Face
2likes284downloads
Model Card

ReDimNet2-B6 Core ML Speaker Embeddings

ReDimNet2-B6 produces local speaker embeddings for comparing clean voice samples. It does not diarize audio or assign names by itself.

Model

PropertyValue
Parameters12.3 million
FormatCompiled Core ML; FP32 frontend/head, FP16 backbone
Compiled size28.6 MiB
Input96,000 mono Float32 samples
Sample rate16 kHz
Window6 seconds
Output192-dimensional L2-normalized embedding
Minimum deploymentmacOS 15 / iOS 18

The checkpoint was trained on VoxBlink2 and VoxCeleb2. The fixed six-second shape avoids the slow Core ML fallback observed with a flexible waveform shape. Applications should repeat clean two-to-six-second speech to fill the input and center-crop longer samples.

Export revision frontend-fp32-v1 preserves waveform normalization, spectral and mel computation, log/feature normalization, pooling and output projection in FP32. The learned backbone remains FP16. This prevents numerical overflow in preprocessing without changing the checkpoint or input/output contract.

Files

FileSizeDescription
ReDimNet2B6.mlmodelc/28.6 MiBPrecompiled Core ML model
config.json<16 KiBContract, export revision, compiled-file SHA-256 and numerical validation
checksums.json<2 KiBSHA-256 of every other published artifact
README.md<8 KiBThis model card
LICENSE1.0 KiBMIT license from the upstream implementation

Performance

Measured on an Apple M5 Pro after two warm-up predictions:

MeasurementResultMeaning
Warm six-second inference13.6 msOne voice-profile embedding
Warm throughput73.6 embeddings/sRepeated six-second windows after warm-up
Numerical validation20/20 configurations and controlsFive deterministic controls across four allowed-device settings

Numerical validation checks finite, normalized outputs and cosine >= 0.999 against the original FP32 model. Controls cover a noisy two-tone waveform, sparse burst, quiet waveform, repeated 600 ms waveform and silence. Allowed-device settings do not identify actual per-operation placement. These checks do not measure diarization, recognition accuracy, or speaker-verification error rate. Thresholds must be calibrated for the intended microphones, languages, and acoustic conditions. Speaker embeddings are useful for labeling; they are not biometric authentication and do not protect against voice spoofing.

Python usage

python
import coremltools as ct
import numpy as np

model = ct.models.CompiledMLModel("ReDimNet2B6.mlmodelc")
audio = np.zeros((1, 96_000), dtype=np.float32)
embedding = model.predict({"audio": audio})["embedding"]

speech-swift

bash
speech embed-speaker voice.wav --engine redimnet2 --json
swift
import SpeechVAD

let model = try await ReDimNet2SpeakerModel.fromPretrained()
let embedding = try model.embed(audio: samples, sampleRate: 16_000)

Source

Converted from the official PalabraAI/ReDimNet2 B6 vb2+vox2_v0 large-margin checkpoint. The source revision and checkpoint SHA-256 are recorded in config.json.

Links