mazesmazes/tiny-audio
Tiny Audio
A tiny, hackable, open-source speech LLM.
It transcribes English with punctuation, capitalization, word timestamps, and speaker labels: 1.8% WER on LibriSpeech test-clean and 7.4% across 12 benchmarks (11,822 samples pooled). Built and trained with Tiny Audio.
[Try the live demo](https://huggingface.co/spaces/mazesmazes/tiny-audio) · [Train your own](https://github.com/alexkroman/tiny-audio) · [Free 3.5-hour course](https://github.com/alexkroman/tiny-audio/blob/main/docs/course/0-course-overview.md)
Quick Start
pip install "transformers>=5.0" peft torch torchaudio librosafrom transformers import pipeline
pipe = pipeline(
"automatic-speech-recognition", model="mazesmazes/tiny-audio", trust_remote_code=True
)
print(pipe("audio.wav")["text"])
# The quarterly revenue grew by 12% according to Dr. Smith.The input can be a file path, a URL, or a 16 kHz float32 numpy array. No post-processing is needed: the model writes punctuation, capitalization, and numbers itself.
Benchmarks
Word error rate (%, lower is better) on 11,822 samples (up to 1,000 per dataset), scored with the Tiny Audio eval harness (ta eval) after text normalization.
† Held out: no data from this source was used in training.
To reproduce a row, or score AssemblyAI, Deepgram, ElevenLabs, or Apple's on-device recognizer on the same samples:
git clone https://github.com/alexkroman/tiny-audio.git && cd tiny-audio && poetry install
poetry run ta eval -m mazesmazes/tiny-audio -d loquacious -n 100
ASSEMBLYAI_API_KEY=... poetry run ta eval -m assemblyai -d loquacious -n 100More Than Plain Text
Batches and GPU
import torch
pipe = pipeline(
"automatic-speech-recognition",
model="mazesmazes/tiny-audio",
trust_remote_code=True,
device="cuda",
torch_dtype=torch.bfloat16,
)
for r in pipe(["audio1.wav", "audio2.wav", "audio3.wav"], batch_size=4):
print(r["text"])Word-level timestamps
return_timestamps=True times every word with Qwen3-ForcedAligner-0.6B and returns them under a words key. Audio of any length works: the pipeline transcribes it in 8-18 s chunks cut at pauses (the model trained on clips up to 19 s), aligns each chunk against its own transcript in batches, and places every word on the recording's timeline.
result = pipe("audio.wav", return_timestamps=True)
print(result["text"])
for word in result["words"]:
print(word)
# {'word': 'Hi,', 'start': 0.0, 'end': 0.24}Speaker diarization
Pass return_speakers=True to label every word with who said it. It also turns on word timestamps. It works on recordings of any length, keeps speaker labels consistent across hour-long meetings, and handles up to 8 speakers. Each speaker is transcribed separately on audio where everyone else is silenced, so overlapping speech is only partly handled.
Diarization needs transformers installed from main until the next release:
pip install git+https://github.com/huggingface/transformersresult = pipe("meeting.wav", return_speakers=True)
print(result["text"])
# Speaker turns (may overlap)
for seg in result["speaker_segments"]:
print(f"{seg['start']:6.2f}-{seg['end']:6.2f} {seg['speaker']}")
# 0.00- 2.90 SPEAKER_0
# 3.36- 6.47 SPEAKER_1
# Words with timestamps and speakers
for w in result["words"]:
print(f"{w['start']:6.2f}-{w['end']:6.2f} {w['speaker']} {w['word']}")
# 0.00- 0.19 SPEAKER_0 Hi,
# 0.22- 0.69 SPEAKER_0 Daniel.Speakers are numbered by when they first speak. The number of speakers is detected automatically. If you know it, pass num_speakers (exact) or max_speakers (an upper bound); voices beyond the cap are not transcribed:
result = pipe("call.wav", return_speakers=True, num_speakers=2)If alignment or diarization fails (for example, on a transformers release without diarization support), the transcript is still returned and the error is reported under result["timestamp_error"] or result["diarization_error"].
Train Your Own
Everything that produced this model is open: the code, the data mix, and the recipe. Swap the encoder, the LLM, or the projector from config and train on your own data. A smoke test runs on a laptop in about five minutes:
git clone https://github.com/alexkroman/tiny-audio.git && cd tiny-audio && poetry install
poetry run python scripts/train.py +experiments=mps_smokeThe free course walks through how the pieces fit, training a model, evaluating it, and publishing it with a demo like this one.
Model Specifications
Training Details
The training recipe is `configs/experiments/granite_qwen_frozen.yaml`.
Limitations
- English only: Not trained on other languages.
- Sample rate: Expects 16kHz audio (other rates are resampled automatically).
- Audio length: A plain
pipe(audio)call decodes the clip in one pass and works best up to about 19 seconds (the training length). For longer audio passreturn_timestamps=Trueorreturn_speakers=True, which transcribe in 8-18 s chunks automatically. - Speaker diarization: At most 8 speakers per recording. The count is detected automatically;
num_speakersandmax_speakerscan cap it, butmin_speakersis not supported (passing it raises aValueError). Needs transformersmainuntil the next release. - Accuracy: May degrade on:
- Far-field and overlapping speech (see AMI SDM)
- Noisy or low-quality audio
- Rare names and domain-specific terminology
Files
Only the projector and LoRA weights are stored here. The encoder (Granite Speech) and decoder (Qwen3.5-2B) are downloaded from their own Hugging Face repos.
The previous GLM-ASR + Qwen3-0.6B model is still available at revision glm-asr-qwen3-0.6b:
pipe = pipeline(
"automatic-speech-recognition",
model="mazesmazes/tiny-audio",
revision="glm-asr-qwen3-0.6b",
trust_remote_code=True,
)Citation
If you use this model, please cite:
@misc{tinyaudio2024,
author = {Alex Kroman},
title = {Tiny Audio: Minimal ASR Training},
year = {2024},
publisher = {GitHub},
url = {https://github.com/alexkroman/tiny-audio}
}Acknowledgments
- Granite Speech for the audio encoder
- Qwen3.5 for the language model
- Qwen3-ForcedAligner for word timestamps
- Nemotron-3-Diarization for speaker diarization
- The LibriHeavy, People's Speech, Common Voice, GigaSpeech, SPGISpeech, VoxPopuli, AMI, and TED-LIUM teams for training data
License
MIT
