Team Ai
Modelpublic

AsadIsmail/whisper-medium-ternary

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
1likes
Model Card

Whisper-medium — Ternary Quantized (tritplane3)

Ternary-quantized version of openai/whisper-medium.

Specifications

PropertyValue
Base Modelopenai/whisper-medium
Parameters769M
Quantizationtritplane3 (240 decoder layers)
Audio encoderFP16 (preserved)
Stored size453 MB
FP16 size~3.1 GB
Compression1.30×

Usage

python
from ternary_quant.inference import load_ternary_model
import torch, numpy as np

model, proc = load_ternary_model("AsadIsmail/whisper-medium-ternary", runtime_mode="cached", device="cpu")
model = model.float()  # Required for encoder compat

# Transcribe audio
import soundfile as sf
audio, sr = sf.read("audio.flac")
inputs = proc(audio.astype(np.float32), sampling_rate=sr, return_tensors="pt")
inputs = {k: v.float() for k, v in inputs.items()}
with torch.no_grad():
    ids = model.generate(**inputs, max_new_tokens=100)
print(proc.batch_decode(ids, skip_special_tokens=True)[0])

Collection

Part of ternary-models.