bjnortier/coreai-whisper-medium-kv-static-float16
Whisper medium — Core AI export (float16, KV cache, fixed-shape decoder)
An Apple Core AI export of OpenAI's whisper-medium, for on-device speech recognition on Apple silicon.
These are not PyTorch weights. The repository holds a single archive of .aimodel assets that run through Apple's Core AI runtime on macOS 27 / iOS 27 and later. They will not load with transformers.
Contents
whisper-medium-kv-static_float16.zip (~1.41 GB) unpacks to a bundle directory:
Configuration
- Architecture: Whisper encoder/decoder
- Precision: float16 · Cross-attention KV: packed · Decoder positions: static
- d_model: 1024 · Decoder layers: 24 · Attention heads: 16
- Vocab: 51865 · Max target positions: 448 · Mel bins: 80
- Encoder window: 30 s (audio is padded to the window)
- Audio in: 16 kHz mono float32
Fixed-shape decoder
This is the -kv-static variant of `bjnortier/coreai-whisper-medium-kv-float16`. The tokenizer and generation config are identical; the decoder fixes every input shape so the GPU can specialize it once instead of on every token. It takes position_ids as [1, 1] holding the current position, writes the new K/V at that position, and attends over all 448 cache slots under a mask (unwritten slots are zeroed, so an uninitialized cache stays safe).
It is an experiment for running the decoder on the GPU. On large-v3-turbo it was no faster than the dynamic decoder on CPU on a Mac and slower on an iPhone 15 Pro; deeper decoders like this one are unmeasured. If you just want transcription, use the dynamic bundle.
Running it requires the `bjnortier` fork of CoreAISpeech: its WhisperDecoder recognises the fixed position_ids shape and sends the position alone. Both kinds of bundle run through the same API.
Language
The weights are the full multilingual Whisper medium — this export is not English-only. What is English is one line of generation_config.json:
"forced_decoder_ids": [[0, 50258], [1, 50259], ...] // 50259 = <|en|>That token is the decoder prefix's language slot. It matters more than it looks: fed French audio, an <|en|> prefix does not fail, it transcribes the audio as if it were English and returns a fluent English translation. Measured on a French clip, that scores ~96% WER against a French reference while emitting a perfectly readable English sentence — a wrong language is silent unless you check for it.
Upstream `apple/coreai-models` has no API for changing the prefix, so with stock CoreAISpeech this bundle transcribes English only. The `bjnortier` fork adds language selection, and with it the same unmodified bundle handles every language Whisper knows:
// Detect the language from the audio (the default).
let (text, _) = try await model.transcribe(pcm: pcm, language: .detect)
// Or name it.
let (text, _) = try await model.transcribe(pcm: pcm, language: .code("fr"))
// Or keep the prefix this bundle shipped with.
let (text, _) = try await model.transcribe(pcm: pcm, language: .bundleDefault)Detection costs one decoder step — Whisper's first prediction after <|startoftranscript|> is the language token. For long-form audio the language is detected once, on the first window, and reused.
No re-export is needed for other languages; the assets here already carry every language token in added_tokens.json.
The prefix also pins <|notimestamps|>. Timestamps still require a different generation_config.json.
Usage
With CirceKit:
import CirceKit
let transcriber = CirceFileTranscriber(
backend: .coreAI(.bundle(bundleURL)),
locale: Locale(identifier: "en_US")
)
try await transcriber.prepare()
let result = try await transcriber.transcribe(fileAt: audioURL)
print(result.text)Or directly with CoreAISpeech from apple/coreai-models:
let model = try await SpeechRecognitionModel(resourcesAt: bundleURL)
let (text, stats) = try await model.transcribe(audioURL: audioURL)On the fork, CirceFileTranscriber passes its locale through as the decode language, so the example above transcribes French simply by asking for a French locale.
License and attribution
Released under the Apache 2.0 licence, the licence of the source model. The original Whisper medium is by OpenAI; this repository only changes its serialization format. Refer to the upstream model card for training data, evaluation results, and intended use.
