Team Ai
Modelpublic

coder543/parakeet-v3-coreai

sourceHugging Facecc-by-4.0updated 7d agoView on Hugging Face
0likes
Model Card

Parakeet TDT 0.6B V3 — Core AI

Apple Core AI conversion of NVIDIA Parakeet V3, pinned to 541d1f99c6b0c3cd0b11a95167540bb8edefd82b. Converted weights retain the upstream Creative Commons Attribution 4.0 license. See LICENSE and NOTICE. This conversion is not endorsed by NVIDIA.

Choose one self-contained bundle:

  • —fast/: full finite Fourier relative attention, 192 encoder positions, fixed 15-second chunks, and an FP16 ANE decoder batching up to 128 independent chunks.
  • —quality/: full sinusoidal relative attention, shared weights for 576 and 3,008 encoder positions (45-second and four-minute inputs), exact tiled convolution subsampling, fixed 240-second chunks, and an FP32 CPU decoder batching up to four independent chunks. Select the smallest shape that fits each chunk; short recordings use the 45-second shape.

Both use W8A16 encoder weights with selected sensitive projections retained in FP16, and four concurrent encoder requests. Neither truncates attention within its chunk or drops input audio. metadata.json supplies available input shapes; runtime.json records the qualified scheduling policy. Weights are shared across static functions in each asset. The labels describe context and execution policy, not a universal quality ranking: longer context did not improve WER in the English recording measured below.

Requires physical Apple silicon on macOS 27 or iOS 27 and a model-specific host runtime. Graphs accept model features and recurrent states, not audio files. The host must implement the frontend, greedy TDT loop and tokenizer described by the metadata and sidecars. Vocabulary: 8,192 nonblank tokens, blank ID 8,192, durations 0/1/2/3/4, and at most ten symbols per frame. The FP16 decoder returns partition/local token-ID pairs so IDs above 2,048 are reconstructed without rounding. Recurrent state is independent between chunks.

Quality host contract: subsampling_output_scale is 16. Multiply the subsampling graph's output by this factor on the host before passing it into the encoder. This power-of-two scaling protects the FP16 projection from overflow; feeding the scaled output directly into the encoder produces incorrect results. Both graph functions use this contract. Fast defaults to scale 1.

TDT emission frames and predicted durations support token/word timings at an 80 ms frame step. These are native alignments, not forced alignment.

Measured performance

M3 MacBook Air (16 GB), macOS 27 build 26A428. Medians of three runs after warmup; frontend, encoder, host rescaling and decoding are included. Preparation, file I/O and chunk planning are excluded. Audio is JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the 18-minute 15-second recording.

BundleAudio durationTranscriptionAudio / elapsed
fast20 seconds0.129 s155.5×
fast18 min 15 s1.942 s563.9×
quality20 seconds0.189 s106.0×
quality18 min 15 s7.973 s137.4×

Tokens and native word timings were stable across runs, with nominal thermal state. The full recording uses 74 fast chunks or five quality chunks. These contexts and execution policies differ.

Whisper-normalized WER against the supplied long-recording reference:

Bundle / referenceWord errorsWER
fast79/22203.56%
FP32 source, same 15-second cuts77/22203.47%
quality81/22203.65%
FP32 source, same four-minute cuts96/22204.32%

With the benchmark's simpler normalization, fast scores 95/2219 and its source 94/2219; quality scores 91/2219 and its source 107/2219. This is one English recording, not a general accuracy ranking. The reference transcript is from stt-bench-matrix. The numerical oracle uses the independently loaded upstream FP32 weights.

Quality's two shapes produce identical valid outputs on the 20-second numerical fixture, with 2.52% encoder relative RMS error versus FP32. Fast's 15.35-second fixture measures 2.99%. All frontend, subsampling and encoder intermediates were finite in separate full-recording checks. Both bundles exactly reproduce the source tokens on additional German and Ukrainian samples; their encoder errors range from 2.33% to 3.60%. These checks do not establish accuracy across every supported language. Quantization and floating-point operation order can change token decisions, including a proper name in the quality English fixture.

Word timings are ordered, bounded and text-preserving; human word-boundary accuracy was not measured. Decoder tests cover all 8,193 token IDs and recurrent FP16-versus-FP32 steps, including odd IDs above the FP16 exact-integer range.

Bounded full-recording traces contain ANE predictions within every encoder and subsampling call: 148 for fast and ten for quality. Fast's batched decoder loop contains 94 further ANE predictions. Neither trace contains target-process GPU intervals. Quality decoding uses the CPU. Compiler manifests mark the used encoder/subsampling and batched decoder functions fully placed on ANE. This is placement evidence, not arithmetic utilization. Quality's sampled peak client-plus-attributed-neural memory is about 1.29 GB, excluding unattributed compiler, driver and system memory.

Comparison with Apple's Core AI Parakeet V3 example

Measured September 28, 2026 on the same M3 MacBook Air (16 GB, macOS 27 build 26A428), using Apple's unmodified `coreai-models` implementation at e7b24da. Both conversions use the same NVIDIA source revision listed above.

At matched 15-second chunk boundaries, this repository's Fast bundle completed the full Moon speech 6.5× faster, with identical Whisper-normalized WER. The five-second Apple default scored worse on this recording; that is a comparison of default configurations, not evidence of a generally more accurate underlying model.

Implementation / configuration20-second excerptComplete speechComplete-speech RTFxComplete-speech WER
Apple default: FP32, 5-second shape0.342 s18.653 s58.7×11.13%
Apple: FP16, 15-second shape0.336 s13.361 s82.0×3.56%
This repository: Fast, W8A16, 15-second shape0.243 s2.048 s534.8×3.56%
This repository: Quality, W8A16, 45/240-second shapes0.187 s7.852 s139.5×3.65%

These are fresh comparison runs, distinct from the qualification measurements above. Each entry is the median of three complete passes after warmup. Timing includes frontend, encoder, decoding and transcript construction; it excludes preparation, file decoding and chunk planning. All implementations process the complete 1,095.32-second recording. Fast uses 74 chunks; Quality uses five. Apple's five-second configuration uses 220 chunks. Apple's runs were at nominal thermal state; Fast's full-recording runs transitioned from nominal to fair. The short-clip Fast result is slower than its earlier qualification measurement; these measured results are reported without substituting the earlier best run.

Apple's static offline API pads or truncates PCM to its exported window. The comparison caller therefore divides the entire recording into consecutive, non-overlapping windows and invokes the original API sequentially, resetting its decoder between windows. Without that caller-side chunking, passing the complete recording to the five-second static export would drop nearly all its audio. Apple was exported with the original script and its declared dependencies: default flags for FP32/5s, or --dtype float16 --audio-seconds 15 for FP16/15s. Both Swift hosts were built in Release configuration.

WER uses the same reference linked above and Whisper English normalization: Apple 5s = 247/2,220 errors; Apple 15s = 79/2,220; Fast = 79/2,220; Quality = 81/2,220. With the benchmark's simpler normalization the corresponding counts are 288/2,219, 96/2,219, 95/2,219 and 91/2,219. The one-error advantage under that normalization is too small to establish an accuracy improvement. All three repetitions of each configuration returned identical text. This is one English recording; it does not establish a multilingual accuracy ranking. Changing Apple's default to the tested variant changes both precision and chunk length, so their individual effects are not isolated.

Runtime and feature tradeoffs

AreaThese bundles and qualified host runtimeApple's tested example
Encoder storageW8A16 with sensitive projections retained in FP16FP32 by default; FP16 optional
Offline schedulingFour concurrent encoders; Fast batches the ANE decoder across up to 128 independent chunksSequential caller-side chunks; scalar predictor and joint calls
Long recordingsBounded chunk processing that retains all input audioStatic offline call is window-limited; caller must split longer audio
Word timingsNative TDT token/word timings with global chunk offsetsPublic offline API returns text and decode statistics, without word timings
Context choicesFast 15s; Quality shares weights across 45s and 240s functionsOne configurable static duration, or a symbolic-length encoder export
StreamingThese bundles are qualified for final transcriptionSeparate buffered-window streaming export and runtime are available

The host features require a compatible runtime; downloading .aimodel files alone does not implement chunking or word timing. The matched-length speedup includes graph design, quantization and scheduling differences, not just an accelerator comparison. Apple's loader uses Core AI's default device selection; its actual placement was not traced in this experiment. The ANE placement qualification for these bundles is documented above.

Apple's dynamic encoder may avoid fixed-window padding, but its README warns that the FP32 dynamic GPU path is unreliable at many shapes. Dynamic and streaming exports were inspected, not benchmarked here. Our Quality configuration offers longer offline context at a throughput and specialization cost; its name does not imply better WER on every recording.

Measured iPhone performance

iPhone 15 Pro Max, iOS 27 build 24A437, Release runtime. Three-run medians after warmup, using the same Moon speech. Includes bounded WAV reading/conversion, chunk planning, frontend, encoder, decoding and output delivery. Preparation, warmup and report writing are excluded. Hardware traces were captured separately.

BundleAudioTranscriptionAudio / elapsed
fast20 seconds0.162 s123.1×
fast18 min 15 s2.498 s438.5×
quality20 seconds0.212 s94.4×
quality18 min 15 s8.767 s124.9×
BundleObserved preparation with specializationCached preparation
fast46.6 s0.622 s
quality489.9 s0.089 s

Preparation includes loading and any specialization requested by Core AI. Underlying driver cache reuse is opaque; these are observed histories, not controlled fresh-install measurements or evidence of an empty system cache. The quality inference runs follow a cooldown after compilation. All reported passes remained foregrounded; nominal and fair thermal states were observed.

Fast's complete recording trace contains 242 unique ANE predictions across 74 encoder calls, 74 subsampling calls and the batched decoder loop, with zero target-process GPU intervals. Fast token IDs and transcript text match the qualified Mac outputs for both input lengths.

Quality's encoder and subsampling calls contain ANE predictions for both shapes, with zero target-process GPU intervals. Quality decoding uses the CPU.

Quality's short output matches the Mac. Full-recording tokens differ but are stable across all phone runs: 65/2,220 errors (2.93% Whisper-normalized WER), versus 81/2,220 (3.65%) on the Mac and 96/2,220 (4.32%) for the matching FP32 source. With the benchmark's simpler normalization, the phone scores 76/2,219 and the Mac 91/2,219. Phone testing here covers this English recording; the additional German and Ukrainian source checks above were run on Mac. This single-recording difference does not establish a general device or quality ranking.

Placement traces establish execution, not ALU utilization or energy use. Concurrent scopes can overlap the same hardware event; aggregate counts use unique intervals.

Device specialization and loading

Source .aimodel files specialize on the device. On the same Mac:

BundleObserved preparation with specializationSubsequent cached preparation
fast34 s0.10–0.21 s
qualityencoder 5 min 32 s; subsampling 32 s0.02–0.05 s

The quality components were measured in separate processes: their specialization times sum to about 6 min 4 s, not a measured single cold whole-bundle load. The encoder measurement uses the exact distributed graph hash. Fast's observed 34-second preparation also has prior cache history. Neither is a controlled fresh-install measurement. Active OS caches were preserved.

Preparation includes Core AI model initialization and function loading; it excludes host sidecar reads, audio I/O, chunk planning, warmup and transcription. These are observed cache histories, not guaranteed device load times. The underlying ANE cache state is not fully observable, and OS/device/application changes can require specialization again.

AoT compiled artifacts are omitted because no significant load-time benefit has been demonstrated. Authoring debug locations were removed while preserving graph signatures and operation counts. Hugging Face provides file checksums for each repository snapshot.