mrfakename/parakeet-redux-webgpu
0
In-browser speech recognition with moondream/parakeet-redux (149M, 1.58-bit encoder) and moondream/parakeet-ultra (627M).
- ONNX Runtime Web: the Conformer encoder runs on WebGPU (WASM fallback), the TDT prediction/joint network runs on WASM, and the token-and-duration greedy loop is plain JS.
- Weights: mrfakename/parakeet-redux-ONNX, mrfakename/parakeet-ultra-ONNX. Downloaded with parallel range requests and kept in the browser Cache API.
- Word timestamps come from the transducer's predicted token durations (80 ms frames).
- Audio longer than 30 s is split at the quietest point between 20 and 30 s.
Example clips: JFK inaugural address (public domain), LibriSpeech and Multilingual LibriSpeech (CC BY 4.0).
