datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.star-wars-dataset
Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21)
One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of
15 titles with a forced-aligned time span, a character label and provenance. No audio or video is
included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips
from your own copies with the pipeline code (export_asr.py, export_diarization.py).
The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.skylar-dataset
Skylar Dataset
A curated speech dataset built for automatic speech recognition (ASR)
benchmarking. It is assembled by streaming samples from existing public
audio datasets and keeping only the ones that pass a fixed set of quality
and diversity rules, organized by target language and audio duration.
This is not a single-source dataset: samples are pulled from multiple
upstream datasets into one unified structure. Re-running the pipeline with
a different source (same or different… See the full description on the dataset page: https://huggingface.co/datasets/skylar-ai-hf/skylar-dataset.tr-full-dataset
TR-Full_dataset
This is a merged speech dataset containing 41427 audio segments from 88 source datasets.
Dataset Information
Total Segments: 41427
Speakers: 222
Languages: tr
Emotions: neutral, angry, sad, happy
Original Datasets: 88
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.AI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.hindi-english-codeswitch-dataset
Hindi-English Code-Switch ASR Transcripts
Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr.
This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here.
Credits
Speech data collection and curation credit: SPRING Lab, IIT Madras.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Emilia-Dataset.stage1a_smoke_data
stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh)
Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP →
frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend:
WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT
encoder — no raw-audio decoding at train time.
113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.SPS-Bopha-Voice-Dataset-v1
VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1
This dataset is formatted for fine-tuning VibeVoice.
Structure
training_data.jsonl: The main manifest file containing transcriptions and paths.
chunks_staging/: Directory containing the audio clips.
Usage with VibeVoice
Clone this repository:
git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1
cd SPS-Bopha-Voice-Dataset-v1
Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.Luciole-Audio-Training-Dataset
Luciole Audio Training Dataset
Dataset description
Luciole Audio Training Dataset is a large, multilingual, multi-task collection of
audio–text conversations used to train the OpenLLM-France
Luciole audio-language models. It adapts a text LLM to understand audio by pairing speech,
music and environmental sounds with instruction-style dialogues (transcription, translation,
spoken question answering, audio/music/sound captioning and question answering, speaker and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Training-Dataset.Luciole-Audio-Evaluation-Dataset
Luciole Audio Evaluation Dataset
Dataset description
Luciole Audio Evaluation Dataset is the held-out evaluation collection used to
benchmark the OpenLLM-France Luciole
audio-language models. It is the evaluation counterpart of the
Luciole Audio Training Dataset:
a multilingual, multi-task set of audio–text conversations, restricted to the test
splits of each source dataset, covering transcription, speech translation, spoken and
music question answering, music… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Evaluation-Dataset.
