Team Ai
Datasetpublic

Squadstack/conversational-streaming-asr-benchmark

SquadStack Conversational Streaming ASR Benchmark (8 kHz) Version 1.0.0 · maintained by SquadStack Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use Key takeaways What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.

sourceHugging Faceotherupdated 1d agoView on Hugging Face
1likes49downloads
Dataset Card

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/01-at-a-glance-dark.png"> <img alt="At a glance: 863 real telesales calls, 11 ASR systems, 5 business domains, a 3x3 grid of code-mixing and noise, 8 kHz telephony." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/01-at-a-glance-light.png"> </picture>

SquadStack Conversational Streaming ASR Benchmark (8 kHz)

Version 1.0.0 · maintained by SquadStack

Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use

Key takeaways

What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on what the customer meant and how fast the transcript arrives?

What to expect from it.

  • —Raw WER overstates the problem. It is 27–55% across the 11 systems. Semantic WER, which counts only errors that would change what the agent does, is 9–22%. For the median system, 32% of word errors matter.
  • —The top four are close. Soniox v5 (8.77%), Arth V2 (8.85%), Arth V1 (9.17%) and Nova 3 (9.31%) sit within about half a point of each other.
  • —Noise costs more than code-mixing. Clean to noisy calls adds 4–12 points; low to high code-mixing adds 2–5.

Semantic WER

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/08-leaderboard-semantic-wer-dark.png"> <img alt="Semantic WER by system for 11 systems on 863 calls, ordered by score, with SquadStack's in-house systems highlighted." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/08-leaderboard-semantic-wer-light.png"> </picture>

Semantic WER, lower is better, pooled over about 67,000 reference words; ★ marks SquadStack's in-house systems. Solid bar = errors that change meaning; pale bar = all word errors (raw WER). Full tables in [LEADERBOARD.md](docs/LEADERBOARD.md).

Latency against accuracy

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/09-latency-accuracy-pareto-dark.png"> <img alt="80th-percentile final-transcript latency (FTR P80, ms) against semantic WER for ten systems on the locked 50-call set." width="960" src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/09-latency-accuracy-pareto-light.png"> </picture>

Semantic WER against FTR P80: the 80th percentile, over all turns, of the time from our end-of-turn finalize signal to the vendor's final reply for that turn. Latency is from the locked 50-call set streamed in real time from a Delhi server; semantic WER is from the 863-call run. Chirp 3 does not support a client-side finalize, so it has no FTR. Method in [LATENCY.md](docs/LATENCY.md), table in [LEADERBOARD.md](docs/LEADERBOARD.md).

It is an evaluation set. It is not for training, and training on it is a licence violation.


What this measures — and what it does not

It measuresIt does not measure
Transcription of spontaneous Hinglish over 8 kHz telephonyBatch transcription quality on clean, wide-band audio
Behaviour on unsegmented audio that is about 80% non-speechPer-utterance accuracy on pre-segmented clips
Whether errors change what a downstream agent would doLiteral string fidelity as an end in itself
80th-percentile time from our end-of-turn signal to the final transcript (50-call set)Throughput, cost at scale, or GPU efficiency
Robustness across five Indian consumer-sales verticalsAny non-telesales domain: medical, legal, broadcast, meetings
Hindi and English in a code-mixed registerThe other 20 scheduled Indian languages
Customer (lead-side) audio onlyAgent-side audio, or full-duplex diarization

If your question is "which streaming STT should run inside my Indian voice agent", this benchmark is built for you. If your question is "which ASR model is best", it is not: use it alongside the Open ASR Leaderboard, FLEURS and IndicVoices, not instead of them.


At a glance

Calls863
Audio29.58 h total · 5.53 h of customer speech (18.7%)
Format8 kHz, mono, PCM16, customer channel extracted from stereo telephony
LanguageHindi–English code-mixed: 16% low, 53% mid, 31% high code-mixing
DomainsMarketplace, Logistics, Travel, BFSI, Education
Campaigns19
Gender16.9% female / 83.1% male (human-annotated)
ReferencesHuman-annotated, turn-separated, code-switched (Hindi in Devanagari, English in Latin)
Splittest only. There is no train split, by design.
Systems benchmarked11: SquadStack in-house and commercial APIs
Headline metricSemantic WER (pooled), with raw WER, its edit types, error classes and P80 latency alongside

Dataset structure

Two tables, both 863 rows, one row per call, joined on sample_id. Loading the dataset without a config name gives you test.

python
from datasets import load_dataset

bench = load_dataset("Squadstack/conversational-streaming-asr-benchmark", "test")["test"]
res   = load_dataset("Squadstack/conversational-streaming-asr-benchmark", "results")["test"]
ConfigRowsColumnsWhat it is
test86311the benchmark: audio, both versions of the reference, and the call's acoustic and demographic profile
results86337the same audio, what the 11 systems returned and every error each one made, as a structured per-error profile

The split is deliberate. test is the thing being released and does not change when a system is added or re-run; results is a snapshot of eleven systems and will. A new baseline never touches the benchmark itself.

Recomputing our numbers

Every accuracy figure in this card is a count over the error profiles in results, and both tables ship:

python
prof = res["soniox_v5_error_profile"]              # one list per call
N    = sum(res["ref_words"])

errs = [e for call in prof for e in call]
sem  = [e for e in errs if e["semantic_check"] == "Yes"]

raw_wer      = len(errs) / N * 100
semantic_wer = len(sem)  / N * 100

If your numbers disagree with ours, we treat it as our bug and log it as an erratum. Latency is measured while audio streams and comes from the run record, which the maintainers hold.

Fields

<details> <summary><b><code>test</code>: 11 fields, one row per call</b></summary>

FieldTypeDescription
sample_idintSerial number, 1 to 863. Stable: cite this, never a row index. Joins to results.
audioAudio(8 kHz)Customer-channel clip, unsegmented: ringing, hold and silence retained. Personal information is replaced by a 1 kHz beep.
duration_secfloat32Whole call, silence included
speech_duration_secfloat32Customer speech only (voice-activity detection)
business_domainstringBFSI / Travel / Marketplace / Logistics / Education
speaker_genderstringmale / female, human-annotated
noise_bandstringclean / moderate / noisy, from DNSMOS bak and NISQA noi on identical 9-second pure-speech chunks
cmi_bandstringlow_mix / mid_mix / high_mix: code-mixing index of the normalised human reference, banded at 0.15 and 0.35
original_referencestringHuman transcript before normalisation, turn-separated. Shipped so you can disagree with our normalisation and redo it from source. Personal information is masked as *****.
normalised_referencestringThe reference every system is scored against. Cleaned once, then locked; one line per turn. Personal information is masked as *****.
ref_wordsint32Words in normalised_reference, counted before masking

</details>

<details> <summary><b><code>results</code>: 37 fields, one row per call</b></summary>

Same 863 rows, joined on sample_id. Three columns per system, so adding a twelfth system adds three columns and changes nothing else.

FieldTypeDescription
sample_idintJoin key to test
audioAudio(8 kHz)The same redacted clip as in test, repeated so the viewer plays it next to the transcripts
normalised_referencestringRepeated from test so this table is scoreable on its own
ref_wordsint32Reference word count, counted before masking
<system>_rawstringWhat the vendor returned, with its own punctuation and digits; personal information is masked as *****
<system>_normalisedstringAfter normalisation and alignment to the reference turns, one line per reference turn (an empty line where the system returned nothing for that turn); the text that was scored, with personal information masked as *****
<system>_error_profilelist&lt;struct&gt;Every error that system made on that call, one record each; words involved in masked text appear as *****

Each error record: word (बीस → बिस for a substitution, the token alone for a deletion or insertion), op (S · D · I), turn_id, semantic_check (Yes · No: would this change what the agent does next), and category (number · negation · name · commitment · content · filler · spelling).

<details> <summary>The 11 systems, in column order</summary>

PrefixSystemVendor
arth_v1Arth V1SquadStack
arth_v2Arth V2SquadStack
nova_3Nova 3Deepgram
scribe_v2Scribe v2 RealtimeElevenLabs
saaras_v4Saaras v4Sarvam
soniox_v5Soniox v5Soniox
fluxFluxDeepgram
cartesia_ink_2Cartesia ink-2Cartesia
gemini_3_5Gemini 3.5 TranscribeGoogle
chirp_3Chirp 3Google
openai_gpt_liveOpenAI gpt-live-transcribeOpenAI

</details>

</details>

Redaction

Personal information (the lead's identity and SquadStack's customer names) is removed everywhere. In the audio it is a 1 kHz beep. In every text field (original_reference, normalised_reference, the system _raw and _normalised texts and the word of each error) it is masked as *****, and the same words are masked everywhere within a call. A mask can also cover a word next to the personal information.

The numbers in this card were computed before masking, and every error record is kept, so counting the error profiles reproduces them exactly. Scoring the masked texts yourself gives slightly different values: the masked words match each other and the reference has about 900 fewer words, so each system's raw and semantic WER come out about 0.3 to 0.5 points lower, with the same ranking. call_sid and any other call identifier are not published; sample_id is a serial number.


How the 863 were chosen

Zero contamination with Arth's training data. Every one of the 863 calls was checked against the audio used to pre-train and fine-tune SquadStack's Arth models. No recording and no speaker in this benchmark appears in either training set.

Every stage is a fixed rule, not a hand pick: positive-outcome calls only (the customer engaged), calls profiled for speech share, noise, code-mixing, vocabulary diversity and content, and a selection that fills a 3 × 3 grid of code-mixing against noise with no domain above a quarter of the speech hours. Personal information and client names are removed from the audio and the transcripts.


Diversity

A benchmark that concentrates on one or two verticals cannot tell you whether a system will hold up on a vertical you have not tested, so the selection is spread deliberately across code-mixing density, acoustic condition and business vertical. Gender is reported, not balanced.

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/03-cmi-noise-grid-dark.png"> <img alt="Panels: code-mixing by noise grid, domain mix, call length, customer speech per call, speaking rate and vocabulary." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/03-cmi-noise-grid-light.png"> </picture>

Stratification: code-mixing × acoustic condition

Nine cells, all populated, from 9 calls (low-mix clean) to 265 (mid-mix moderate). The grid is unbalanced, and the marginal slices inherit that. Read the nine cells, not the three margins. Code-mixing is 16.1% low, 53.3% mid and 30.6% high; noise is 12.4% clean, 59.0% moderate and 28.6% noisy.

Domain

Logistics 260 calls, Marketplace 240, Travel 187, BFSI 122, Education 54, across 19 campaigns. No vertical exceeds 23% of the speech hours, and between 23% and 40% of each vertical's vocabulary appears in no other vertical, which is what makes the domain slices worth reading separately.

Speech versus silence

Four-fifths of this audio is not speech: ringing, hold, pauses and the agent's turn. This is the single biggest structural difference from clip-based read-speech sets, and it is deliberate. The honest size of this benchmark is 5.53 hours of speech, not 29.58 hours of audio, and that is the number quoted everywhere in this card.

Vocabulary

67,443 tokens and 4,086 distinct words. 40.7% of words appear exactly once and Yule's K is 89.8: conversational speech repeats a small set of function words, so the diagnostic numbers are the ones describing the tail. That shape carries a risk: a large, easy filler bucket can mask failures on rare domain words, which is why meaning-changing errors and their classes are reported next to the aggregate.

Speakers and conditions

146 calls are female speakers (16.9%), from human annotation, and 717 are male. The median call is 80 seconds (P10 38, P90 284) carrying 14 seconds of customer speech (4–51), spoken at a median 188 words per minute (118–276). Caller accent is not annotated.


Why we score meaning

Both of these turns contain exactly one wrong word.

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/17-two-errors-dark.png"> <img alt="Two reference turns with one wrong word each. 'तीन हजार लगभग' heard as 'मतलब तीन हजार लगभग' keeps the amount intact. 'छबीस' (26) heard as 'छत्तीस' (36) changes the number the customer said." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/17-two-errors-light.png"> </picture>

ReferenceSystem outputWord errorsJudge's verdictWhat happens next
तीन हजार लगभग (about three thousand)मतलब तीन हजार लगभग1 (one extra word)harmless · fillerNothing. The amount is right.
छबीस (twenty-six)छत्तीस (thirty-six)1 (one wrong word)counts · numberThe agent hears the wrong number.

Raw word error rate counts one error in each, so it scores them alike. A voice agent does not read the transcript, it acts on it. So the headline metric here is semantic WER: errors that could change the meaning of a turn in a way that affects the downstream task, divided by reference words. In everyday telesales speech, fillers and repetitions dominate the raw error count and rarely matter. For each error, the judge sees the reference turn, the system output and the error, answers would the agent act differently?, and tags a class (here filler and number).


Evaluation protocol

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/12-evaluation-pipeline-dark.png"> <img alt="Scoring flow: reference normalisation, vendor normalisation and alignment to reference turns, deterministic error counting per turn, LLM judging of each error, semantic WER." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/12-evaluation-pipeline-light.png"> </picture>

Three things about this protocol are load-bearing.

The same recordings and the same pipeline for every system. All 11 systems transcribed the same 863 calls, and one pinned model (Gemini 3.7 Flash, temperature 0) runs every LLM step with the same prompts. If a system returned nothing for a call, every reference word of that call counts as deleted.

The normalisers are ours, not yours. Submissions carry raw transcripts and both normalisation passes run on our side. If each vendor normalised its own output, the leaderboard would measure normalisation effort as much as recognition quality. Numeric values, negation, named entities and word order are never normalised: a wrong number is an error, a differently formatted number is not. हाँ, 15। becomes हां fifteen; रिक्वायरमेंट becomes requirement.

Each error is judged on its own. One error, one question (would the agent act differently?), answered inside the turn it occurred in, with no sight of the running total. Every error is then tagged with a class, and both the verdict and the class ship in results, so the judgement is auditable rather than asserted.

Aligned per turn, not per call

Vendors cut speech differently and sometimes drop a whole turn. Aligned over the whole call, a matcher can pair a word with its neighbour in the next turn. In one real call the reference says "fifty percent" and the system says "sixty percent"; whole-call alignment pairs हां (the previous turn) with "sixty" and marks "fifty" deleted, while per-turn alignment keeps it as one number error in the turn where it happened.

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/18-turn-split-dark.png"> <img alt="A vendor returning ten lines for a six-turn reference, aligned back to the reference turns." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/18-turn-split-light.png"> </picture>

On a 240-call subset, per-turn counting was 0.5 to 1.0 point stricter in token WER for five of the six systems tested, because words can no longer be matched across turns.

The metrics

MetricDefinition
Semantic WER (headline)meaning-changing errors ÷ reference words. A voice agent does not read the transcript, it acts on it.
Raw WER(S+D+I)/N, with the three components published separately so the filtering is auditable
Error classeach meaning-changing error tagged number, negation, name, commitment or content; the classes sum to semantic WER
Latency (FTR)time from our end-of-turn finalize signal to the vendor's final reply for that turn, 80th percentile (P80) over all turns, on the locked 50-call set; TTFS adds a constant 480 ms (method)

Latency measurement

Each call is streamed in real time; at every end of a customer turn, decided by our production VAD (stop 0.5 s, an effective 0.48 s at 8 kHz), we send the system its finalize signal. FTR is the time from that signal to the system's reply for the turn, on a locked 50-call set at 5 concurrency, one system at a time, from one server in Delhi. TTFS = FTR + 480 ms: the time from actual speech end, as in pipecat-ai/stt-benchmark, which includes the VAD wait (Pipecat's benchmark defaults to a 0.2 s stop, so compare FTR, not TTFS, across the two). Full method, P50–P99 per system and the Soniox region comparison: `docs/LATENCY.md`.

Planned for the next release: confidence intervals and paired comparisons, character error rate, entity WER (errors divided by entity tokens), hallucination on silence, and a measured agreement between the judge and human adjudication.


Baseline results

Ordered by semantic WER, lowest first. SquadStack's in-house systems are marked ★ and ranked with the others: comparing a model tuned on this kind of audio with a general-purpose endpoint as peers is the usual way vendor benchmarks mislead.

SystemVendorCategorySemantic WERRaw WERSubDelInsFTR P80 (ms)
Soniox v5SonioxCommercial8.7727.2912.665.459.1789.4
Arth V2 ★SquadStackIn-house8.8529.1413.178.057.9267.0
Arth V1 ★SquadStackIn-house9.1729.8313.428.967.4567.0
Nova 3DeepgramCommercial9.3130.9811.6914.494.8093.8
Saaras v4SarvamCommercial10.8433.2514.736.2612.26176.7
Cartesia ink-2CartesiaCommercial11.5836.2714.693.8417.73183.6
Gemini 3.5 TranscribeGoogleCommercial11.9335.8214.4814.676.68264.2
OpenAI gpt-live-transcribeOpenAICommercial12.3038.1418.1613.526.46830.8
Chirp 3GoogleCommercial12.4734.0315.1410.967.93n/a
FluxDeepgramCommercial14.3337.7618.2812.916.56124.4
Scribe v2 RealtimeElevenLabsCommercial22.0954.9523.999.0921.8786.5

Pooled over about 67,000 reference words (67,388–67,393 per system); Sub, Del and Ins are percentages of reference words and add to raw WER. Latency: FTR P80 on the locked 50-call set (see above). Point estimates: systems within about a point of each other should not be ranked on this table alone.

What raw WER hides

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/14-error-composition-dark.png"> <img alt="For each system, raw WER split into the part that changes meaning and the part that does not, and into substitutions, deletions and insertions." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/14-error-composition-light.png"> </picture>

Between 30% and 40% of every system's raw word errors would change what an agent does; for the median system it is 32%. A 29.1% raw WER is an 8.8% rate of errors that matter, so a decision made on raw WER reads a gap roughly three times its real size.

The ordering mostly survives, and the exceptions are in the middle of the table. Chirp 3 is sixth on raw WER and ninth on semantic WER: 37% of its errors change meaning, against 30–32% for the four leading systems. Cartesia ink-2 moves up two places because about half of its raw errors are insertions that the judge mostly treats as harmless. OpenAI also moves up two, passing Flux and Chirp 3, which lose a larger share of their errors to meaning. The top five keep their places.

What the errors land on

Each meaning-changing error carries a class, and the classes sum to each system's semantic WER (percent of reference words):

SystemSemantic WERNumberNegationNameCommitmentContent
Soniox v58.771.070.430.770.306.18
Arth V2 ★8.851.100.520.920.425.89
Arth V1 ★9.171.110.520.990.416.14
Nova 39.311.240.551.110.485.92
Saaras v410.841.290.590.900.507.55
Cartesia ink-211.581.150.600.930.428.47
Gemini 3.5 Transcribe11.931.430.790.950.658.08
OpenAI gpt-live-transcribe12.301.590.681.020.468.55
Chirp 312.471.910.700.910.528.41
Flux14.331.680.861.360.649.79
Scribe v2 Realtime22.092.111.331.470.9216.25

Two systems with the same semantic WER are not equally usable. Number and negation errors change the outcome of a call (a wrong amount is quoted, a decline is heard as a yes) while a content error usually costs a retry. Content dominates for every system; number errors run from 1.1% to 2.1% of reference words.

By slice

<picture> <source media="(prefers-color-scheme: dark)" srcset="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/11-semantic-wer-by-slice-dark.png"> <img alt="Heatmap of semantic WER for every system across noise, code-mixing, domain and gender slices." src="https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark/resolve/main/assets/11-semantic-wer-by-slice-light.png"> </picture>

  • —Noise is the largest factor. Semantic WER rises 4–12 points from clean to noisy calls (Soniox 4.9 → 12.7, Cartesia 5.6 → 17.1).
  • —Code-mixing costs less than noise. Going from low to high mixing adds 2–5 points.
  • —Marketplace is the hardest vertical, Education the easiest. Soniox: 13.9% on Marketplace, 6.5% on Education (54 calls).
  • —No consistent gender gap (female 8.1–12.7% vs male 8.8–14.7% for all but Scribe v2), on 146 female calls.

Uses

Direct use

Comparing speech-to-text systems for deployment in Indian consumer-sales voice agents. The claim this benchmark licenses you to make is narrow and specific:

On Hindi–English telesales calls over 8 kHz telephony, system A made meaning-changing transcription errors at rate X and returned a final transcript within Y ms of the end-of-turn signal for 80% of turns.

It also supports per-condition diagnosis: which systems degrade on heavy code-mixing or noise, which lose numbers, which flip negations.

Out-of-scope use

  • —Training or fine-tuning any model. This is a licence violation, not a recommendation: see Licensing.
  • —Speaker identification, speaker verification, voice cloning, or any attempt to re-identify a caller.
  • —Claims about ASR quality in general, in other languages, in other domains, or on clean wide-band audio.
  • —Claims about a vertical represented by few calls, such as Education (54 calls).
  • —Procurement decisions on latency alone. Latency here is from one server location and a 50-call set; reproduce the measurement from your own network location before acting on it.

Licensing

Four separate statements, because they genuinely differ:

ComponentLicence
AudioBenchmark Evaluation Licence 1.0: evaluation only, no redistribution, no training, no speaker identification. Recordings remain the property of SquadStack and its clients.
Reference transcripts and annotationsCC BY-NC 4.0, attribution to SquadStack
Compilation, metadata and schemaCC BY 4.0
Evaluation code (harness and metrics, supplied on request)Apache 2.0

Full text in `LICENSE.md`; the access terms you accept at the gate are in `TERMS_OF_USE.md`. Commercial evaluation is permitted; publishing results is encouraged; training is not permitted under any of the four.


Citation

bibtex
@dataset{telesales_asr_benchmark_2026,
  title        = {SquadStack Conversational Streaming ASR Benchmark (8 kHz)},
  author       = {{SquadStack MLOps Team}},
  year         = {2026},
  version      = {1.0.0},
  publisher    = {SquadStack},
  url          = {https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark},
  note         = {Evaluation-only benchmark of 863 real Indian telesales calls,
                  8 kHz narrowband, Hindi-English code-mixed}
}

SquadStack MLOps Team (2026). SquadStack Conversational Streaming ASR Benchmark (8 kHz) (Version 1.0.0) [Data set]. SquadStack.


Acknowledgements

Curation methodology builds on the code-mixing index of Gambäck & Das (2016), MATTR (Covington & McFall, 2010), DNSMOS (Reddy et al., Microsoft) and NISQA (Mittag et al., TU Berlin). The evaluation design draws on the Open ASR Leaderboard, and the semantic-WER framing and final-transcript latency clock of pipecat-ai/stt-benchmark.

Dataset card authors SquadStack MLOps Team