Arm/whisper-large-v3-quantized.w4a8
Whisper Large v3 W4A8
Model Overview
- Model architecture: Whisper large-v3
- Input: Audio
- Output: Text
- Model optimizations:
- Weight quantization: INT4
- Activation quantization: INT8
- Base model: `openai/whisper-large-v3`
This model is a W4A8 GPTQ-quantized version of `openai/whisper-large-v3`. It is intended for automatic speech recognition evaluation and inference on Arm CPU systems using vLLM.
The quantization workflow uses `llmcompressor` to quantize Whisper linear layers with 4-bit integer weights and 8-bit integer input activations. The language-model projection head is left unquantized. The model was quantized without additional fine-tuning.
Intended Use
This model is intended for:
- Automatic speech recognition with Whisper-compatible tooling.
- Experiments comparing BF16 and compressed Whisper inference.
- Arm CPU vLLM deployments where reduced model size is useful.
This model is not intended to improve transcription quality over the base model. Users should validate quality, latency, memory use, and supported runtime behavior on their target hardware and workload before deployment.
Deployment
Use with vLLM
from vllm import LLM, SamplingParams
from vllm.assets.audio import AudioAsset
llm = LLM(
model="Arm/whisper-large-v3-quantized.w4a8",
max_model_len=448,
max_num_seqs=400,
limit_mm_per_prompt={"audio": 1},
)
inputs = {
"encoder_prompt": {
"prompt": "",
"multi_modal_data": {
"audio": AudioAsset("winning_call").audio_and_sample_rate,
},
},
"decoder_prompt": "<|startoftranscript|>",
}
outputs = llm.generate(inputs, SamplingParams(temperature=0.0, max_tokens=128))
print(outputs[0].outputs[0].text)Quantization Details
Evaluation
Evaluation was run with `lmms-eval` using the whisper_vllm model interface on LibriSpeech and FLEURS. Lower WER is better.
On average our INT4 implementation is able to recover 97.2% of BF16 WER.
Reproduce Quantization
Create a fresh quantization environment:
python -m venv .quantize
source .quantize/bin/activate
pip install llmcompressor
pip install torchcodec --index-url https://download.pytorch.org/whl/cpuRun quantization:
python quantize.py \
--model_path openai/whisper-large-v3 \
--save_dir data/model_dir \
--num_calibration_samples 1024The output is written to:
data/model_dir/whisper-large-v3-quantized.w4a8Reproduce Evaluation
Create a fresh evaluation environment:
python -m venv .eval
source .eval/bin/activate
export VLLM_VERSION=0.23.0
pip install "https://github.com/vllm-project/vllm/releases/download/v${VLLM_VERSION}/vllm-${VLLM_VERSION}+cpu-cp38-abi3-manylinux_2_34_aarch64.whl" --extra-index-url https://download.pytorch.org/whl/cpu
pip install editdistance
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu
git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git ~/lmms-eval
pip install -e ~/lmms-evalRun the W4A8 evaluation:
lmms-eval \
--model whisper_vllm \
--model_args "pretrained=Arm/whisper-large-v3-quantized.w4a8" \
--tasks librispeech_test_other,librispeech_test_clean,fleurs \
--batch_size 64 \
--output_path results/w4a8/full_suiteLimitations
- Reported evaluations cover LibriSpeech and two FLEURS language splits only.
- The calibration samples use English transcription examples.
- Runtime support depends on vLLM, compressed-tensors, and target hardware.
- Quantization can change outputs, especially on languages, accents, domains, audio conditions, and decoding settings not covered by the reported evaluation.
Ethical Considerations
This model inherits the capabilities and limitations of Whisper large-v3. ASR systems can produce incorrect transcripts and may perform unevenly across languages, accents, dialects, speakers, domains, and recording conditions. Do not use transcripts as the sole basis for high-stakes decisions without human review.
About this version
This repository contains a W4A8 quantized version of OpenAI’s Whisper large-v3 model. Arm quantized the model using INT4 weight quantization and INT8 activation quantization to enable more efficient execution with vLLM on Arm-based platforms. No additional training or fine-tuning was applied by Arm. The original model architecture, intended automatic speech recognition and speech translation use cases, and known limitations remain applicable, although quantization may affect numerical behavior and accuracy.
Original model and documentation
For full details of the original model, please refer to the original OpenAI Whisper large-v3 model card: https://huggingface.co/openai/whisper-large-v3
Purpose of this release
Arm provides this quantized model to enable developers to evaluate and build applications using W4A8 Whisper large-v3 inference with vLLM on Arm-based systems. Users should validate its accuracy and behavior under the languages, audio conditions, decoding settings, and deployment environment relevant to their application.
