TPekarekRosin/modicol
MoDiCoL - A Modular Diagnostic Continual Learning Dataset for ASR MoDiCoL is a speech dataset designed to study the robustness of ASR models to different drift factors in a controlled, continual setting. We construct MoDiCoL using a systematic factorial design that enables a rigorous evaluation of linguistic, speaker, and acoustic variation with clearly defined experimental runs. By combining real-world and synthetic speech with a configuration-dependent augmentation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TPekarekRosin/modicol.
MoDiCoL - A Modular Diagnostic Continual Learning Dataset for ASR
MoDiCoL is a speech dataset designed to study the robustness of ASR models to different drift factors in a controlled, continual setting. We construct MoDiCoL using a systematic factorial design that enables a rigorous evaluation of linguistic, speaker, and acoustic variation with clearly defined experimental runs. By combining real-world and synthetic speech with a configuration-dependent augmentation pipeline, MoDiCoL allows a targeted evaluation of how different types of factorial drift affect model adaptation and forgetting.
Dataset Overview
MoDiCoL covers three factor categories: Linguistic Content, Speaker Characteristics, and Acoustic Environment.
- Linguistic Content:
- vocabularydomain: medical, airtraffic_control, other
- speech_style: read, spontaneous, conversational
- Speaker Characteristics:
- age: child, adult, elderly
- accent: english, seasian, other
- health: healthy, impaired
- pauses: yes, no
- disfluencies: yes, no
- Acoustic Environment:
- noise_type: clean, babble, fan
- snr_level: clean, 10dB, 20dB
- distance: close, far
We created the dataset with a L27 orthogonal array (OA) and two foldover dimensions to accommodate the six three-dimensional and four two-dimensional factors. With cyclic permutation, we created 108 distinct run configuration (run001 - run108). Each of these runs contains 75 samples, which leads to 8100 samples overall.
The dataset is compiled from multiple open-source datasets, and supplemented with synthetic speech to fill less common or non-existent factor-level combinations. Each sample is marked with its source, please refere to the paper for a comprehensive list.
Overall, MoDiCoL contains 18.79 hours of speech, of which 14.08 are synthetic, with an average sample length of 8.35 seconds.
Data Fields:
Each row contains:
audio: waveform file (16 kHz audio)text: transcription of speech- factor levels: age, speechstyle, snrlevel, etc.
f0: pitch level descriptorduration: audio duration in secondssynthetic: whether sample is synthetic or real-world- augmentation flags: noise, pauses, disfluencies, etc.
source: dataset origin or DOI
Usage
Load a specific run
from datasets import load_dataset
ds = load_dataset("TPekarekRosin/modicol", "run_001")
sample = ds["train"][0]
print(sample["text"])
print(sample["audio"])
Load the Continual Learning Curriculum
The paper defines a CL curriculum to study robustness transfer across drift types.
from datasets import load_dataset
# task 1: Acoustic Environment Drift
ds = load_dataset("TPekarekRosin/modicol", "t1_environment")
# task 2: Speaker Characteristics Drift
ds = load_dataset("TPekarekRosin/modicol", "t2_speaker")
# task 3: Linguistic Content Drift
ds = load_dataset("TPekarekRosin/modicol", "t3_linguistic")
# task 4: Compound Drift
ds = load_dataset("TPekarekRosin/modicol", "t4_compound")
Load the entire dataset
from datasets import load_dataset
ds = load_dataset("TPekarekRosin/modicol", "all")Audio Format
- Sample rate: 16,000 Hz
- Format: WAV
- Channels: mono
Data Augmentation
Step 1 is performed for all files, while steps 2 to 6 are done according to the specifications of the respective run configuration.
- Denoising. Background noise is suppressed for all samples.
- Disfluency Insertion. We add synthetic filler words with voice cloning to match the speaker. These filler words are inserted the same way as pauses are (Step 4).
- Impairment Simulation. We introduce impairment via prosodic and spectral modifications.
- Pause Insertion/Removal. Pauses are inserted by detecting existing silent segments or by randomly selecting insertion points, or noticeable silent segments are removed instead.
- Distance Simulation. We add reverberation to model a greater distance between the speaker and the microphone.
- Noise Injection. Background noise is added according to the predefined signal-to-noise ratio (SNR) and noise type.
Citation
MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition was accepted at Interspeech 2026.
If you use this dataset, please use the following citation:
@inproceedings{pekarekrosin2026_modicol,
title={MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition},
author={Theresa Pekarek Rosin and Matthias Kerzel and Stefan Wermter},
year={2026},
booktitle={Interspeech 2026},
doi={https://doi.org/10.48550/arXiv.2606.14459}
}
ToDo
- [ ] Add augmentation pipeline script.
- [ ] Add more information about dataset metadata.
- [ ] Update citation once the paper is published.
