ERISLab/LisTAya-transcripts
LisTAya transcripts: the test-set evaluations of the LisTAya study This dataset holds the reference and the model output for every utterance of every test-set evaluation in the paper Does Regional Decoder Specialization Help Low-Resource ASR Based on the SLAM-ASR Framework? (ROCLING 2026). The trained models are described in the model card ERISLab/LisTAya and listed in the collection… See the full description on the dataset page: https://huggingface.co/datasets/ERISLab/LisTAya-transcripts.
LisTAya transcripts: the test-set evaluations of the LisTAya study
This dataset holds the reference and the model output for every utterance of every test-set evaluation in the paper Does Regional Decoder Specialization Help Low-Resource ASR Based on the SLAM-ASR Framework? (ROCLING 2026). The trained models are described in the model card ERISLab/LisTAya and listed in the collection https://huggingface.co/collections/ERISLab/listaya-slam-asr-projectors-on-tiny-aya-rocling-2026-6ab031414095f3df8805ed7b.
Configs
Each config has one split, test, with one row per utterance per run. A run is one model evaluated on one test set.
Test sets
There are 43 test sets, all test splits:
- out-of-domain: the FLEURS test split (
google/fleurs) of 11 languages; FLEURS does not cover Kreol Seselwa. - in-domain: the WorldSpeech test split of the dialect each language trains on, one per language (12). Hausa's is
ha_td, because its checkpoints were selected on theha_ngtest split. Kreol Seselwa's is thetest_cleansplit ofERISLab/WorldSpeechconfigcrs_sc; every other language's comes fromdisco-eth/WorldSpeech. - trained dialect: the WorldSpeech test split of a second dialect in the training data:
sw_tz,ur_in(2). - held-out dialect: the WorldSpeech test splits of dialects absent from training (18):
en_au,en_jm,en_ke,en_nz,en_pk,en_sl,en_zm,es_ar,es_cl,es_co,es_es,es_pe,es_pr,es_py,es_uy,fr_cd,fr_ci,ta_lk.
Every baseline is evaluated on all 43 test sets, except that mistralai/Voxtral-Mini-3B-2507 and openai/whisper-medium produced no output on Kreol Seselwa, so baselines holds 127 runs (3 x 43 - 2).
Columns
Text normalisation
The references are the transcription column of FLEURS and the human_transcript column of WorldSpeech. References and model outputs pass through the same normaliser before scoring, for every language: lowercasing; removal of text inside square brackets, angle brackets and parentheses; Unicode NFKD decomposition, with nonspacing combining marks (category Mn) deleted and all other marks, symbols and punctuation replaced by a space; removal of any remaining character that is neither a word character nor whitespace; and collapsing of whitespace. In scripts that write vowels as combining signs, such as Devanagari and Tamil, the normalised text therefore keeps only part of each syllable. Utterances whose normalised reference is empty are left out of the evaluation, so sample_index equals the row index of the source split only where none were left out.
Relation to the paper
The CER of a run is the corpus-level character error rate of its rows in percent, computed with the cer metric of the evaluate library and rounded to two decimals. Recomputed from these rows, it equals the CER the paper uses for all 299 runs. The paper's per-test-set tables print, for each test set, the lowest of these CERs among the three baselines and among the language's four trained checkpoints: the in-domain and out-of-domain test sets in the table that compares training with the baselines, and the trained and held-out dialects in the table of dialects.
Load the transcripts and score a run
import evaluate
from datasets import load_dataset
rows = load_dataset("ERISLab/LisTAya-transcripts", "listaya", split="test")
run = rows.filter(lambda r: r["model"] == "ERISLab/q2a_openai_whisper-medium_CohereLabs_tiny-aya-global_ws-en_us-500" and r["eval_dataset"] == "google/fleurs")
cer = evaluate.load("cer").compute(references=run["reference"], predictions=run["hypothesis"])
print(f"{100 * cer:.2f}") # 6.22Licence and attribution
The dataset is released under CC BY-NC 4.0. The references are normalised transcripts from FLEURS (google/fleurs, CC BY 4.0) and WorldSpeech (disco-eth/WorldSpeech, CC BY-NC 4.0), whose non-commercial term this release follows. The Kreol Seselwa references come from ERISLab/WorldSpeech, config crs_sc, recorded sessions of the National Assembly of Seychelles.
Citation
@inproceedings{rios-etal-2026-regional,
title = "Does Regional Decoder Specialization Help Low-Resource {ASR} Based on the {SLAM}-{ASR} Framework?",
author = "Rios, Edwin Arkel and
Zaruma, Jocelyn and
Ewoorkar, Girish and
Sourabh, Sneh and
Mack, Julian and
Juan, Hung-Hui and
Huang, Stephen and
Lai, Bo-Cheng",
booktitle = "Proceedings of the 38th Conference on Computational Linguistics and Speech Processing (ROCLING 2026)",
year = "2026",
publisher = "Association for Computational Linguistics"
}