Team Ai
Modelpublic

bjnortier/coreai-whisper-medium-kv-static-float16

sourceHugging Faceapache-2.0updated 1d agoView on Hugging Face
0likes
Model Card

Whisper medium — Core AI export (float16, KV cache, fixed-shape decoder)

An Apple Core AI export of OpenAI's whisper-medium, for on-device speech recognition on Apple silicon.

These are not PyTorch weights. The repository holds a single archive of .aimodel assets that run through Apple's Core AI runtime on macOS 27 / iOS 27 and later. They will not load with transformers.

Contents

whisper-medium-kv-static_float16.zip (~1.41 GB) unpacks to a bundle directory:

FilePurpose
metadata.jsonBundle descriptor: architecture and model dimensions
encoder.aimodelAudio encoder
decoder.aimodelFixed-shape text decoder with packed cross-attention KV cache
generation_config.jsonForced decoder prefix, EOS, max new tokens
tokenizer.json, tokenizer_config.jsonTokenizer files
added_tokens.jsonSpecial tokens (language, task) → ids

Configuration

  • —Architecture: Whisper encoder/decoder
  • —Precision: float16 · Cross-attention KV: packed · Decoder positions: static
  • —d_model: 1024 · Decoder layers: 24 · Attention heads: 16
  • —Vocab: 51865 · Max target positions: 448 · Mel bins: 80
  • —Encoder window: 30 s (audio is padded to the window)
  • —Audio in: 16 kHz mono float32

Fixed-shape decoder

This is the -kv-static variant of `bjnortier/coreai-whisper-medium-kv-float16`. The tokenizer and generation config are identical; the decoder fixes every input shape so the GPU can specialize it once instead of on every token. It takes position_ids as [1, 1] holding the current position, writes the new K/V at that position, and attends over all 448 cache slots under a mask (unwritten slots are zeroed, so an uninitialized cache stays safe).

It is an experiment for running the decoder on the GPU. On large-v3-turbo it was no faster than the dynamic decoder on CPU on a Mac and slower on an iPhone 15 Pro; deeper decoders like this one are unmeasured. If you just want transcription, use the dynamic bundle.

Running it requires the `bjnortier` fork of CoreAISpeech: its WhisperDecoder recognises the fixed position_ids shape and sends the position alone. Both kinds of bundle run through the same API.

Language

The weights are the full multilingual Whisper medium — this export is not English-only. What is English is one line of generation_config.json:

json
"forced_decoder_ids": [[0, 50258], [1, 50259], ...]   // 50259 = <|en|>

That token is the decoder prefix's language slot. It matters more than it looks: fed French audio, an <|en|> prefix does not fail, it transcribes the audio as if it were English and returns a fluent English translation. Measured on a French clip, that scores ~96% WER against a French reference while emitting a perfectly readable English sentence — a wrong language is silent unless you check for it.

Upstream `apple/coreai-models` has no API for changing the prefix, so with stock CoreAISpeech this bundle transcribes English only. The `bjnortier` fork adds language selection, and with it the same unmodified bundle handles every language Whisper knows:

swift
// Detect the language from the audio (the default).
let (text, _) = try await model.transcribe(pcm: pcm, language: .detect)

// Or name it.
let (text, _) = try await model.transcribe(pcm: pcm, language: .code("fr"))

// Or keep the prefix this bundle shipped with.
let (text, _) = try await model.transcribe(pcm: pcm, language: .bundleDefault)

Detection costs one decoder step — Whisper's first prediction after <|startoftranscript|> is the language token. For long-form audio the language is detected once, on the first window, and reused.

No re-export is needed for other languages; the assets here already carry every language token in added_tokens.json.

The prefix also pins <|notimestamps|>. Timestamps still require a different generation_config.json.

Usage

With CirceKit:

swift
import CirceKit

let transcriber = CirceFileTranscriber(
    backend: .coreAI(.bundle(bundleURL)),
    locale: Locale(identifier: "en_US")
)
try await transcriber.prepare()
let result = try await transcriber.transcribe(fileAt: audioURL)
print(result.text)

Or directly with CoreAISpeech from apple/coreai-models:

swift
let model = try await SpeechRecognitionModel(resourcesAt: bundleURL)
let (text, stats) = try await model.transcribe(audioURL: audioURL)

On the fork, CirceFileTranscriber passes its locale through as the decode language, so the example above transcribes French simply by asking for a French locale.

License and attribution

Released under the Apache 2.0 licence, the licence of the source model. The original Whisper medium is by OpenAI; this repository only changes its serialization format. Refer to the upstream model card for training data, evaluation results, and intended use.