Team Ai
Datasetpublic

mueller91/MLAAD-tiny

Welcome to MLAAD-tiny MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging. Many teaching environments (e.g. Colab, Kaggle, university notebooks -- se this notebook for example) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD. Download git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/MLAAD-tiny.

sourceHugging Facecc-by-nc-4.0updated 4mo agoView on Hugging Face
3likes2.8kdownloads
Dataset Card

Welcome to MLAAD-tiny

MLAAD-tiny is a very small subset of the full MLAAD dataset, designed for education, prototyping, and debugging.

Many teaching environments (e.g. Colab, Kaggle, university notebooks -- se this notebook for example) impose strict storage limits, which makes large-scale audio deepfake datasets impractical to use. To address this, we provide MLAAD-tiny, a compact yet representative version of MLAAD.

Download

git lfs install
git clone https://huggingface.co/datasets/mueller91/MLAAD-tiny

Dataset composition

Bona-fide

  • —Source: M-AILABS
  • —~6,000 audio files
  • —~1.9 GB
  • —English

Spoof

  • —64 TTS systems
  • —100 samples per system (randomly selected from MLAAD)
  • —~6,400 audio files
  • —~2.3 GB
  • —English (for training) and German (for testing)

License

  • —Bona-fide audio is redistributed from M-AILABS under its original license (see original/LICENSE).
  • —Spoofed audio is redistributed under the MLAAD v8 license (CC BY-NC 4.0).
mueller91/MLAAD-tiny · Team Ai