Squadstack/conversational-streaming-asr-benchmark
SquadStack Conversational Streaming ASR Benchmark (8 kHz) Version 1.0.0 · maintained by SquadStack Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use Key takeaways What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/01-at-a-glance-dark.png"> <img alt="At a glance: 863 real telesales calls, 11 ASR systems, 5 business domains, a 3x3 grid of code-mixing and noise, 8 kHz telephony." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/01-at-a-glance-light.png"> </picture>
SquadStack Conversational Streaming ASR Benchmark (8 kHz)
Version 1.0.0 · maintained by SquadStack
Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use
Key takeaways
What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on what the customer meant and how fast the transcript arrives?
What to expect from it.
- Raw WER overstates the problem. It is 27–55% across the 11 systems. Semantic WER, which counts only errors that would change what the agent does, is 9–22%. For the median system, 32% of word errors matter.
- The top four are close. Soniox v5 (8.77%), Arth V2 (8.85%), Arth V1 (9.17%) and Nova 3 (9.31%) sit within about half a point of each other.
- Noise costs more than code-mixing. Clean to noisy calls adds 4–12 points; low to high code-mixing adds 2–5.
Semantic WER
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/08-leaderboard-semantic-wer-dark.png"> <img alt="Semantic WER by system for 11 systems on 863 calls, ordered by score, with SquadStack's in-house systems highlighted." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/08-leaderboard-semantic-wer-light.png"> </picture>
Semantic WER, lower is better, pooled over about 67,000 reference words; ★ marks SquadStack's in-house systems. Solid bar = errors that change meaning; pale bar = all word errors (raw WER). Full tables in [LEADERBOARD.md](docs/LEADERBOARD.md).
Latency against accuracy
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/09-latency-accuracy-pareto-dark.png"> <img alt="80th-percentile final-transcript latency (FTR P80, ms) against semantic WER for ten systems on the locked 50-call set." width="960" src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/09-latency-accuracy-pareto-light.png"> </picture>
Semantic WER against FTR P80: the 80th percentile, over all turns, of the time from our end-of-turn finalize signal to the vendor's final reply for that turn. Latency is from the locked 50-call set streamed in real time from a Delhi server; semantic WER is from the 863-call run. Chirp 3 does not support a client-side finalize, so it has no FTR. Method in [LATENCY.md](docs/LATENCY.md), table in [LEADERBOARD.md](docs/LEADERBOARD.md).
It is an evaluation set. It is not for training, and training on it is a licence violation.
What this measures — and what it does not
If your question is "which streaming STT should run inside my Indian voice agent", this benchmark is built for you. If your question is "which ASR model is best", it is not: use it alongside the Open ASR Leaderboard, FLEURS and IndicVoices, not instead of them.
At a glance
Dataset structure
Two tables, both 863 rows, one row per call, joined on sample_id. Loading the dataset without a config name gives you test.
from datasets import load_dataset
bench = load_dataset("Squadstack/conversational-streaming-asr-benchmark", "test")["test"]
res = load_dataset("Squadstack/conversational-streaming-asr-benchmark", "results")["test"]The split is deliberate. test is the thing being released and does not change when a system is added or re-run; results is a snapshot of eleven systems and will. A new baseline never touches the benchmark itself.
Recomputing our numbers
Every accuracy figure in this card is a count over the error profiles in results, and both tables ship:
prof = res["soniox_v5_error_profile"] # one list per call
N = sum(res["ref_words"])
errs = [e for call in prof for e in call]
sem = [e for e in errs if e["semantic_check"] == "Yes"]
raw_wer = len(errs) / N * 100
semantic_wer = len(sem) / N * 100If your numbers disagree with ours, we treat it as our bug and log it as an erratum. Latency is measured while audio streams and comes from the run record, which the maintainers hold.
Fields
<details> <summary><b><code>test</code>: 11 fields, one row per call</b></summary>
</details>
<details> <summary><b><code>results</code>: 37 fields, one row per call</b></summary>
Same 863 rows, joined on sample_id. Three columns per system, so adding a twelfth system adds three columns and changes nothing else.
Each error record: word (बीस → बिस for a substitution, the token alone for a deletion or insertion), op (S · D · I), turn_id, semantic_check (Yes · No: would this change what the agent does next), and category (number · negation · name · commitment · content · filler · spelling).
<details> <summary>The 11 systems, in column order</summary>
</details>
</details>
Redaction
Personal information (the lead's identity and SquadStack's customer names) is removed everywhere. In the audio it is a 1 kHz beep. In every text field (original_reference, normalised_reference, the system _raw and _normalised texts and the word of each error) it is masked as *****, and the same words are masked everywhere within a call. A mask can also cover a word next to the personal information.
The numbers in this card were computed before masking, and every error record is kept, so counting the error profiles reproduces them exactly. Scoring the masked texts yourself gives slightly different values: the masked words match each other and the reference has about 900 fewer words, so each system's raw and semantic WER come out about 0.3 to 0.5 points lower, with the same ranking. call_sid and any other call identifier are not published; sample_id is a serial number.
How the 863 were chosen
Zero contamination with Arth's training data. Every one of the 863 calls was checked against the audio used to pre-train and fine-tune SquadStack's Arth models. No recording and no speaker in this benchmark appears in either training set.
Every stage is a fixed rule, not a hand pick: positive-outcome calls only (the customer engaged), calls profiled for speech share, noise, code-mixing, vocabulary diversity and content, and a selection that fills a 3 × 3 grid of code-mixing against noise with no domain above a quarter of the speech hours. Personal information and client names are removed from the audio and the transcripts.
Diversity
A benchmark that concentrates on one or two verticals cannot tell you whether a system will hold up on a vertical you have not tested, so the selection is spread deliberately across code-mixing density, acoustic condition and business vertical. Gender is reported, not balanced.
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/03-cmi-noise-grid-dark.png"> <img alt="Panels: code-mixing by noise grid, domain mix, call length, customer speech per call, speaking rate and vocabulary." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/03-cmi-noise-grid-light.png"> </picture>
Stratification: code-mixing × acoustic condition
Nine cells, all populated, from 9 calls (low-mix clean) to 265 (mid-mix moderate). The grid is unbalanced, and the marginal slices inherit that. Read the nine cells, not the three margins. Code-mixing is 16.1% low, 53.3% mid and 30.6% high; noise is 12.4% clean, 59.0% moderate and 28.6% noisy.
Domain
Logistics 260 calls, Marketplace 240, Travel 187, BFSI 122, Education 54, across 19 campaigns. No vertical exceeds 23% of the speech hours, and between 23% and 40% of each vertical's vocabulary appears in no other vertical, which is what makes the domain slices worth reading separately.
Speech versus silence
Four-fifths of this audio is not speech: ringing, hold, pauses and the agent's turn. This is the single biggest structural difference from clip-based read-speech sets, and it is deliberate. The honest size of this benchmark is 5.53 hours of speech, not 29.58 hours of audio, and that is the number quoted everywhere in this card.
Vocabulary
67,443 tokens and 4,086 distinct words. 40.7% of words appear exactly once and Yule's K is 89.8: conversational speech repeats a small set of function words, so the diagnostic numbers are the ones describing the tail. That shape carries a risk: a large, easy filler bucket can mask failures on rare domain words, which is why meaning-changing errors and their classes are reported next to the aggregate.
Speakers and conditions
146 calls are female speakers (16.9%), from human annotation, and 717 are male. The median call is 80 seconds (P10 38, P90 284) carrying 14 seconds of customer speech (4–51), spoken at a median 188 words per minute (118–276). Caller accent is not annotated.
Why we score meaning
Both of these turns contain exactly one wrong word.
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/17-two-errors-dark.png"> <img alt="Two reference turns with one wrong word each. 'तीन हजार लगभग' heard as 'मतलब तीन हजार लगभग' keeps the amount intact. 'छबीस' (26) heard as 'छत्तीस' (36) changes the number the customer said." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/17-two-errors-light.png"> </picture>
Raw word error rate counts one error in each, so it scores them alike. A voice agent does not read the transcript, it acts on it. So the headline metric here is semantic WER: errors that could change the meaning of a turn in a way that affects the downstream task, divided by reference words. In everyday telesales speech, fillers and repetitions dominate the raw error count and rarely matter. For each error, the judge sees the reference turn, the system output and the error, answers would the agent act differently?, and tags a class (here filler and number).
Evaluation protocol
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/12-evaluation-pipeline-dark.png"> <img alt="Scoring flow: reference normalisation, vendor normalisation and alignment to reference turns, deterministic error counting per turn, LLM judging of each error, semantic WER." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/12-evaluation-pipeline-light.png"> </picture>
Three things about this protocol are load-bearing.
The same recordings and the same pipeline for every system. All 11 systems transcribed the same 863 calls, and one pinned model (Gemini 3.7 Flash, temperature 0) runs every LLM step with the same prompts. If a system returned nothing for a call, every reference word of that call counts as deleted.
The normalisers are ours, not yours. Submissions carry raw transcripts and both normalisation passes run on our side. If each vendor normalised its own output, the leaderboard would measure normalisation effort as much as recognition quality. Numeric values, negation, named entities and word order are never normalised: a wrong number is an error, a differently formatted number is not. हाँ, 15। becomes हां fifteen; रिक्वायरमेंट becomes requirement.
Each error is judged on its own. One error, one question (would the agent act differently?), answered inside the turn it occurred in, with no sight of the running total. Every error is then tagged with a class, and both the verdict and the class ship in results, so the judgement is auditable rather than asserted.
Aligned per turn, not per call
Vendors cut speech differently and sometimes drop a whole turn. Aligned over the whole call, a matcher can pair a word with its neighbour in the next turn. In one real call the reference says "fifty percent" and the system says "sixty percent"; whole-call alignment pairs हां (the previous turn) with "sixty" and marks "fifty" deleted, while per-turn alignment keeps it as one number error in the turn where it happened.
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/18-turn-split-dark.png"> <img alt="A vendor returning ten lines for a six-turn reference, aligned back to the reference turns." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/18-turn-split-light.png"> </picture>
On a 240-call subset, per-turn counting was 0.5 to 1.0 point stricter in token WER for five of the six systems tested, because words can no longer be matched across turns.
The metrics
Latency measurement
Each call is streamed in real time; at every end of a customer turn, decided by our production VAD (stop 0.5 s, an effective 0.48 s at 8 kHz), we send the system its finalize signal. FTR is the time from that signal to the system's reply for the turn, on a locked 50-call set at 5 concurrency, one system at a time, from one server in Delhi. TTFS = FTR + 480 ms: the time from actual speech end, as in pipecat-ai/stt-benchmark, which includes the VAD wait (Pipecat's benchmark defaults to a 0.2 s stop, so compare FTR, not TTFS, across the two). Full method, P50–P99 per system and the Soniox region comparison: `docs/LATENCY.md`.
Planned for the next release: confidence intervals and paired comparisons, character error rate, entity WER (errors divided by entity tokens), hallucination on silence, and a measured agreement between the judge and human adjudication.
Baseline results
Ordered by semantic WER, lowest first. SquadStack's in-house systems are marked ★ and ranked with the others: comparing a model tuned on this kind of audio with a general-purpose endpoint as peers is the usual way vendor benchmarks mislead.
Pooled over about 67,000 reference words (67,388–67,393 per system); Sub, Del and Ins are percentages of reference words and add to raw WER. Latency: FTR P80 on the locked 50-call set (see above). Point estimates: systems within about a point of each other should not be ranked on this table alone.
What raw WER hides
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/14-error-composition-dark.png"> <img alt="For each system, raw WER split into the part that changes meaning and the part that does not, and into substitutions, deletions and insertions." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/14-error-composition-light.png"> </picture>
Between 30% and 40% of every system's raw word errors would change what an agent does; for the median system it is 32%. A 29.1% raw WER is an 8.8% rate of errors that matter, so a decision made on raw WER reads a gap roughly three times its real size.
The ordering mostly survives, and the exceptions are in the middle of the table. Chirp 3 is sixth on raw WER and ninth on semantic WER: 37% of its errors change meaning, against 30–32% for the four leading systems. Cartesia ink-2 moves up two places because about half of its raw errors are insertions that the judge mostly treats as harmless. OpenAI also moves up two, passing Flux and Chirp 3, which lose a larger share of their errors to meaning. The top five keep their places.
What the errors land on
Each meaning-changing error carries a class, and the classes sum to each system's semantic WER (percent of reference words):
Two systems with the same semantic WER are not equally usable. Number and negation errors change the outcome of a call (a wrong amount is quoted, a decline is heard as a yes) while a content error usually costs a retry. Content dominates for every system; number errors run from 1.1% to 2.1% of reference words.
By slice
<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/11-semantic-wer-by-slice-dark.png"> <img alt="Heatmap of semantic WER for every system across noise, code-mixing, domain and gender slices." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/11-semantic-wer-by-slice-light.png"> </picture>
- Noise is the largest factor. Semantic WER rises 4–12 points from clean to noisy calls (Soniox 4.9 → 12.7, Cartesia 5.6 → 17.1).
- Code-mixing costs less than noise. Going from low to high mixing adds 2–5 points.
- Marketplace is the hardest vertical, Education the easiest. Soniox: 13.9% on Marketplace, 6.5% on Education (54 calls).
- No consistent gender gap (female 8.1–12.7% vs male 8.8–14.7% for all but Scribe v2), on 146 female calls.
Uses
Direct use
Comparing speech-to-text systems for deployment in Indian consumer-sales voice agents. The claim this benchmark licenses you to make is narrow and specific:
On Hindi–English telesales calls over 8 kHz telephony, system A made meaning-changing transcription errors at rate X and returned a final transcript within Y ms of the end-of-turn signal for 80% of turns.
It also supports per-condition diagnosis: which systems degrade on heavy code-mixing or noise, which lose numbers, which flip negations.
Out-of-scope use
- Training or fine-tuning any model. This is a licence violation, not a recommendation: see Licensing.
- Speaker identification, speaker verification, voice cloning, or any attempt to re-identify a caller.
- Claims about ASR quality in general, in other languages, in other domains, or on clean wide-band audio.
- Claims about a vertical represented by few calls, such as Education (54 calls).
- Procurement decisions on latency alone. Latency here is from one server location and a 50-call set; reproduce the measurement from your own network location before acting on it.
Licensing
Four separate statements, because they genuinely differ:
Full text in `LICENSE.md`; the access terms you accept at the gate are in `TERMS_OF_USE.md`. Commercial evaluation is permitted; publishing results is encouraged; training is not permitted under any of the four.
Citation
@dataset{telesales_asr_benchmark_2026,
title = {SquadStack Conversational Streaming ASR Benchmark (8 kHz)},
author = {{SquadStack MLOps Team}},
year = {2026},
version = {1.0.0},
publisher = {SquadStack},
url = {https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark},
note = {Evaluation-only benchmark of 863 real Indian telesales calls,
8 kHz narrowband, Hindi-English code-mixed}
}SquadStack MLOps Team (2026). SquadStack Conversational Streaming ASR Benchmark (8 kHz) (Version 1.0.0) [Data set]. SquadStack.
Acknowledgements
Curation methodology builds on the code-mixing index of Gambäck & Das (2016), MATTR (Covington & McFall, 2010), DNSMOS (Reddy et al., Microsoft) and NISQA (Mittag et al., TU Berlin). The evaluation design draws on the Open ASR Leaderboard, and the semantic-WER framing and final-transcript latency clock of pipecat-ai/stt-benchmark.
Dataset card authors SquadStack MLOps Team
