coder543/parakeet-ultra-glade
Parakeet Ultra for Glade
Converted model assets and runtime configuration for Glade.
Glade conversion of Moondream Parakeet Ultra, a post-trained NVIDIA Parakeet TDT 0.6B V3 model. Source revision: 73175eb7aeb0d82f1e2a6b53b3aabc10a90bcd0b. The weights retain the upstream CC-BY-4.0 license. See LICENSE and NOTICE. This conversion is not endorsed by Moondream or NVIDIA.
The repository contains one self-contained bundle. It preserves the model's native speech head to find pauses, selecting chunks of at most 30 seconds. All audio is retained, including silence. The speech head runs on CPU over ANE subsampling features; the acoustic encoder and batched TDT decoder prefer ANE. No separate VAD model is needed.
Requires physical Apple silicon with macOS 27 or iOS 27 and Glade's model-specific host runtime. These graphs accept features and recurrent states, not audio files. The host implements the frontend, chunk planning, greedy TDT loop and tokenizer using the supplied metadata and sidecars.
Runtime configuration
- W8A16 encoder with selected sensitive projections retained in FP16; FP16 batched decoder. Original attention and subsampling masks are preserved.
- One 384-position encoder shape, four concurrent encoder requests and a decoder batch capacity of 128 independent chunks.
- The native speech head uses bounded 120-second normalization blocks. The chunker selects pauses within the 30-second limit.
metadata.jsondescribes graph conventions;runtime.jsonowns chunking and scheduling. Its internalfastprofile selects the ANE decoder; it does not identify a second downloadable variant.- 8,193 vocabulary entries, blank ID 8,192, durations 0/1/2/3/4 and at most ten symbols per acoustic frame. Batched token IDs are represented as an exact FP16 partition/local pair:
partition * 2048 + local.
The model retains V3's 25-language coverage. Native token/word timing uses TDT emission frames and predicted durations at an 80 ms frame step. The host must join subwords and attach punctuation without stretching word spans across silence. These are decoder alignments, not forced alignment.
Measured performance
JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the 18-minute 15-second recording. Three-run warmed medians, excluding model preparation. Measurements were made September 30, 2026.
The Mac timing includes frontend, encoder and decoding; file reading and chunk planning are excluded. Native VAD analysis took another 0.325 s for the full recording, yielding about 2.79 s / 393× including that analysis. Phone timing includes bounded file reading, VAD, inference and text delivery. Phone runs stayed in nominal thermal state and foreground throughout. This is one recording, not a general speed or accuracy ranking.
Peak measured full-recording memory was 1.27 GB on Mac (combined client and separately attributed ANE allocations) and 1.14 GB on iPhone (client footprint, which already includes its ANE allocation).
Preparation includes the runtime's required functions, not inference. The Mac already had related experimental assets cached; its 7.26-second result is not a clean-device compilation estimate. Core AI's underlying cache state is not fully observable, and OS updates or eviction can cause recompilation.
Qualification
- The packaged Mac bundle and all three phone runs returned identical tokens and native-VAD boundaries for the complete recording, consuming every sample.
- A separate complete phone trace recorded 278 ANE prediction intervals and zero app GPU intervals. All 39 encoder and 39 transcription subsampling scopes contained ANE work. Overlapping calls make per-stage hardware counts non-additive; this establishes placement, not ALU utilization.
- On this recording, Whisper-normalized WER was 54/2,220 errors (2.43%), matching the source model's error count under the same frozen cut plan.
- Seven public multilingual regression samples matched source token sequences, with native timing bounds checked. This is conversion coverage, not a comprehensive multilingual accuracy evaluation.
- The quantized speech head is approximate: a 120-second comparison against the FP32 source had 0.00247 mean absolute probability error, with six of 1,500 frames crossing the 0.5 threshold. Native-VAD cuts are not claimed to reproduce every FP32 source boundary.
Download the complete repository snapshot. Source locations, evaluation audio, reference tensors and compiled device caches are excluded. Hugging Face's snapshot metadata supplies file hashes for integrity verification.
Experimental Core ML bundle for Apple Watch
coreml-watch/bundle.tar contains compiled watchOS 27 Core ML graphs, the host frontend and TDT decoder, runtime metadata, license files and loading.json with measured per-file loading estimates. Extract this uncompressed POSIX ustar archive into one model directory. The Core AI bundle remains available separately.
The public Core ML runtime requests CPU and Neural Engine execution with the default specialization strategy. Six encoder stages remain loaded; the frontend and TDT decoder run on CPU. The native speech head selects pauses for chunks up to 30 seconds, retaining all audio including silence. Recording replay uses up to eight independent chunks per decoder batch. Native word timings are available; they are not forced-aligned boundaries.
Measured on Apple Watch Ultra 4, watchOS 27.0.1, Release runtime, resident graphs, one warmup and three uninterrupted foreground runs without real-time pauses. The input is JFK’s “We choose to go to the Moon” speech.
Processing includes audio reads, the frontend, native VAD/chunk planning, encoder, decoder and transcript delivery; it excludes preparation, warmup and result serialization. Preparation is an observed first strategy-specific load with existing caches preserved, not a guaranteed globally cold measurement. Cache reuse is demonstrated across these two launches, not guaranteed after eviction or OS updates.
The full recording’s 39 native VAD cuts match the independent FP32 source. Three Watch runs produce identical tokens, and their transcripts match the source after lowercasing and removing punctuation. This single recording does not establish a general accuracy ranking. Full-recording hardware tracing confirmed ANE predictions for subsampling, native VAD and all six encoder stages. The separate Neural Engine allocation ledger reports about 1 GiB during inference; that accounting does not prove all bytes are physically resident at once.
