Team Ai
Modelpublic

Arm/whisper-large-v3-quantized.w4a8

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes48downloads
Model Card

Whisper Large v3 W4A8

Model Overview

  • —Model architecture: Whisper large-v3
  • —Input: Audio
  • —Output: Text
  • —Model optimizations:
  • —Weight quantization: INT4
  • —Activation quantization: INT8
  • —Base model: `openai/whisper-large-v3`

This model is a W4A8 GPTQ-quantized version of `openai/whisper-large-v3`. It is intended for automatic speech recognition evaluation and inference on Arm CPU systems using vLLM.

The quantization workflow uses `llmcompressor` to quantize Whisper linear layers with 4-bit integer weights and 8-bit integer input activations. The language-model projection head is left unquantized. The model was quantized without additional fine-tuning.

Intended Use

This model is intended for:

  • —Automatic speech recognition with Whisper-compatible tooling.
  • —Experiments comparing BF16 and compressed Whisper inference.
  • —Arm CPU vLLM deployments where reduced model size is useful.

This model is not intended to improve transcription quality over the base model. Users should validate quality, latency, memory use, and supported runtime behavior on their target hardware and workload before deployment.

Deployment

Use with vLLM

python
from vllm import LLM, SamplingParams
from vllm.assets.audio import AudioAsset

llm = LLM(
    model="Arm/whisper-large-v3-quantized.w4a8",
    max_model_len=448,
    max_num_seqs=400,
    limit_mm_per_prompt={"audio": 1},
)

inputs = {
    "encoder_prompt": {
        "prompt": "",
        "multi_modal_data": {
            "audio": AudioAsset("winning_call").audio_and_sample_rate,
        },
    },
    "decoder_prompt": "<|startoftranscript|>",
}

outputs = llm.generate(inputs, SamplingParams(temperature=0.0, max_tokens=128))
print(outputs[0].outputs[0].text)

Quantization Details

FieldValue
Base modelopenai/whisper-large-v3
Quantization methodGPTQ
Weight precisionINT4
Weight strategyPer-channel, symmetric
Input activation precisionINT8
Activation strategyDynamic, per-token, symmetric
Quantized modulesLinear layers
Unquantized modulesproj_out
Calibration datasetMLCommons/peoples_speech, subset test, split test
Calibration task prefixEnglish transcription
Calibration samples used for reported run1,024
Maximum calibration sequence length2,048

Evaluation

Evaluation was run with `lmms-eval` using the whisper_vllm model interface on LibriSpeech and FLEURS. Lower WER is better.

BenchmarkSplitBF16 WERW4A8 WERBF16/W4A8 Recovery
LibiriSpeech (WER)test-clean2.15172.198997.9%
LibiriSpeech (WER)test-other3.93524.086596.3%
Fleurs (WER)cmnhanscn7.79078.174195.3%
Fleurs (WER)en4.04424.078599.2%

On average our INT4 implementation is able to recover 97.2% of BF16 WER.

Reproduce Quantization

Create a fresh quantization environment:

bash
python -m venv .quantize
source .quantize/bin/activate
pip install llmcompressor
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu

Run quantization:

bash
python quantize.py \
  --model_path openai/whisper-large-v3 \
  --save_dir data/model_dir \
  --num_calibration_samples 1024

The output is written to:

text
data/model_dir/whisper-large-v3-quantized.w4a8

Reproduce Evaluation

Create a fresh evaluation environment:

bash
python -m venv .eval
source .eval/bin/activate
export VLLM_VERSION=0.23.0
pip install "https://github.com/vllm-project/vllm/releases/download/v${VLLM_VERSION}/vllm-${VLLM_VERSION}+cpu-cp38-abi3-manylinux_2_34_aarch64.whl" --extra-index-url https://download.pytorch.org/whl/cpu
pip install editdistance
pip install torchcodec --index-url https://download.pytorch.org/whl/cpu

git clone https://github.com/EvolvingLMMs-Lab/lmms-eval.git ~/lmms-eval
pip install -e ~/lmms-eval

Run the W4A8 evaluation:

bash
lmms-eval \
  --model whisper_vllm \
  --model_args "pretrained=Arm/whisper-large-v3-quantized.w4a8" \
  --tasks librispeech_test_other,librispeech_test_clean,fleurs \
  --batch_size 64 \
  --output_path results/w4a8/full_suite

Limitations

  • —Reported evaluations cover LibriSpeech and two FLEURS language splits only.
  • —The calibration samples use English transcription examples.
  • —Runtime support depends on vLLM, compressed-tensors, and target hardware.
  • —Quantization can change outputs, especially on languages, accents, domains, audio conditions, and decoding settings not covered by the reported evaluation.

Ethical Considerations

This model inherits the capabilities and limitations of Whisper large-v3. ASR systems can produce incorrect transcripts and may perform unevenly across languages, accents, dialects, speakers, domains, and recording conditions. Do not use transcripts as the sole basis for high-stakes decisions without human review.

About this version

This repository contains a W4A8 quantized version of OpenAI’s Whisper large-v3 model. Arm quantized the model using INT4 weight quantization and INT8 activation quantization to enable more efficient execution with vLLM on Arm-based platforms. No additional training or fine-tuning was applied by Arm. The original model architecture, intended automatic speech recognition and speech translation use cases, and known limitations remain applicable, although quantization may affect numerical behavior and accuracy.

Original model and documentation

For full details of the original model, please refer to the original OpenAI Whisper large-v3 model card: https://huggingface.co/openai/whisper-large-v3

Purpose of this release

Arm provides this quantized model to enable developers to evaluate and build applications using W4A8 Whisper large-v3 inference with vLLM on Arm-based systems. Users should validate its accuracy and behavior under the languages, audio conditions, decoding settings, and deployment environment relevant to their application.