Team Ai
Datasetpublic

TPekarekRosin/modicol

MoDiCoL - A Modular Diagnostic Continual Learning Dataset for ASR MoDiCoL is a speech dataset designed to study the robustness of ASR models to different drift factors in a controlled, continual setting. We construct MoDiCoL using a systematic factorial design that enables a rigorous evaluation of linguistic, speaker, and acoustic variation with clearly defined experimental runs. By combining real-world and synthetic speech with a configuration-dependent augmentation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/TPekarekRosin/modicol.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes182downloads
Dataset Card

MoDiCoL - A Modular Diagnostic Continual Learning Dataset for ASR

MoDiCoL is a speech dataset designed to study the robustness of ASR models to different drift factors in a controlled, continual setting. We construct MoDiCoL using a systematic factorial design that enables a rigorous evaluation of linguistic, speaker, and acoustic variation with clearly defined experimental runs. By combining real-world and synthetic speech with a configuration-dependent augmentation pipeline, MoDiCoL allows a targeted evaluation of how different types of factorial drift affect model adaptation and forgetting.

Dataset Overview

MoDiCoL covers three factor categories: Linguistic Content, Speaker Characteristics, and Acoustic Environment.

  • —Linguistic Content:
  • —vocabularydomain: medical, airtraffic_control, other
  • —speech_style: read, spontaneous, conversational
  • —Speaker Characteristics:
  • —age: child, adult, elderly
  • —accent: english, seasian, other
  • —health: healthy, impaired
  • —pauses: yes, no
  • —disfluencies: yes, no
  • —Acoustic Environment:
  • —noise_type: clean, babble, fan
  • —snr_level: clean, 10dB, 20dB
  • —distance: close, far

We created the dataset with a L27 orthogonal array (OA) and two foldover dimensions to accommodate the six three-dimensional and four two-dimensional factors. With cyclic permutation, we created 108 distinct run configuration (run001 - run108). Each of these runs contains 75 samples, which leads to 8100 samples overall.

The dataset is compiled from multiple open-source datasets, and supplemented with synthetic speech to fill less common or non-existent factor-level combinations. Each sample is marked with its source, please refere to the paper for a comprehensive list.

Overall, MoDiCoL contains 18.79 hours of speech, of which 14.08 are synthetic, with an average sample length of 8.35 seconds.

Data Fields:

Each row contains:

  • —audio: waveform file (16 kHz audio)
  • —text: transcription of speech
  • —factor levels: age, speechstyle, snrlevel, etc.
  • —f0: pitch level descriptor
  • —duration: audio duration in seconds
  • —synthetic: whether sample is synthetic or real-world
  • —augmentation flags: noise, pauses, disfluencies, etc.
  • —source: dataset origin or DOI

Usage

Load a specific run

python
from datasets import load_dataset

ds = load_dataset("TPekarekRosin/modicol", "run_001")

sample = ds["train"][0]
print(sample["text"])
print(sample["audio"])

Load the Continual Learning Curriculum

The paper defines a CL curriculum to study robustness transfer across drift types.

python
from datasets import load_dataset

# task 1: Acoustic Environment Drift
ds = load_dataset("TPekarekRosin/modicol", "t1_environment")

# task 2: Speaker Characteristics Drift
ds = load_dataset("TPekarekRosin/modicol", "t2_speaker")

# task 3: Linguistic Content Drift
ds = load_dataset("TPekarekRosin/modicol", "t3_linguistic")

# task 4: Compound Drift
ds = load_dataset("TPekarekRosin/modicol", "t4_compound")

Load the entire dataset

python
from datasets import load_dataset

ds = load_dataset("TPekarekRosin/modicol", "all")

Audio Format

  • —Sample rate: 16,000 Hz
  • —Format: WAV
  • —Channels: mono

Data Augmentation

Step 1 is performed for all files, while steps 2 to 6 are done according to the specifications of the respective run configuration.

  1. 1.Denoising. Background noise is suppressed for all samples.
  2. 2.Disfluency Insertion. We add synthetic filler words with voice cloning to match the speaker. These filler words are inserted the same way as pauses are (Step 4).
  3. 3.Impairment Simulation. We introduce impairment via prosodic and spectral modifications.
  4. 4.Pause Insertion/Removal. Pauses are inserted by detecting existing silent segments or by randomly selecting insertion points, or noticeable silent segments are removed instead.
  5. 5.Distance Simulation. We add reverberation to model a greater distance between the speaker and the microphone.
  6. 6.Noise Injection. Background noise is added according to the predefined signal-to-noise ratio (SNR) and noise type.

Citation

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition was accepted at Interspeech 2026.

If you use this dataset, please use the following citation:

@inproceedings{pekarekrosin2026_modicol,
      title={MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition}, 
      author={Theresa Pekarek Rosin and Matthias Kerzel and Stefan Wermter},
      year={2026},
      booktitle={Interspeech 2026},
      doi={https://doi.org/10.48550/arXiv.2606.14459}
}

ToDo

  • —[ ] Add augmentation pipeline script.
  • —[ ] Add more information about dataset metadata.
  • —[ ] Update citation once the paper is published.