Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LEMAS-Project /LEMAS-Dataset-train Overview This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project). It contains a large-scale training set (150k+ hours) and a curated evaluation set (500 utterances per language) covering 10 languages, all with word-level alignment. Fields key: unique utterance identifier; the first two characters indicate the language ID audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.texttext-to-speech100M<n<1B90 likes6.5k downloads6mo agoHugging Face02shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes4.4k downloads9mo agoHugging Face03IMJONEZZ /star-wars-dataset Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with the pipeline code (export_asr.py, export_diarization.py). The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.tabularautomatic-speech-recognition10K<n<100K3 likes345 downloads19d agoHugging Face04skylar-ai-hf /skylar-dataset Skylar Dataset A curated speech dataset built for automatic speech recognition (ASR) benchmarking. It is assembled by streaming samples from existing public audio datasets and keeping only the ones that pass a fixed set of quality and diversity rules, organized by target language and audio duration. This is not a single-source dataset: samples are pulled from multiple upstream datasets into one unified structure. Re-running the pipeline with a different source (same or different… See the full description on the dataset page: https://huggingface.co/datasets/skylar-ai-hf/skylar-dataset.audioautomatic-speech-recognition1K<n<10K0 likes244 downloads10d agoHugging Face05Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes112 downloads1y agoHugging Face06Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes88 downloads2mo agoHugging Face07demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes78 downloads5mo agoHugging Face08oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes73 downloads6mo agoHugging Face09Abhisingh-18 /hindi-english-codeswitch-dataset Hindi-English Code-Switch ASR Transcripts Text transcripts and metadata for a large bilingual Hindi-English code-switch ASR training corpus, used to train Abhisingh-18/hindi-english-codeswitch-asr. This release contains transcripts and metadata only — no audio files. Audio was sourced from multiple corpora and institutions and is not redistributed here. Credits Speech data collection and curation credit: SPRING Lab, IIT Madras. Contents File… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/hindi-english-codeswitch-dataset.textautomatic-speech-recognition1M<n<10M0 likes26 downloads2mo agoHugging Face10samson-ailabs /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Emilia-Dataset.texttext-to-speech100M<n<1B0 likes25 downloads5mo agoHugging Face11Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes21 downloads3mo agoHugging Face12Tnaot /SPS-Bopha-Voice-Dataset-v1gated VibeVoice Fine-Tuning Dataset: SPS-Bopha-Voice-Dataset-v1 This dataset is formatted for fine-tuning VibeVoice. Structure training_data.jsonl: The main manifest file containing transcriptions and paths. chunks_staging/: Directory containing the audio clips. Usage with VibeVoice Clone this repository: git clone https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1 cd SPS-Bopha-Voice-Dataset-v1 Run the training script pointing to… See the full description on the dataset page: https://huggingface.co/datasets/Tnaot/SPS-Bopha-Voice-Dataset-v1.audiotext-to-speech1K<n<10K0 likes9 downloads11mo agoHugging Face13OpenLLM-France /Luciole-Audio-Training-Dataset Luciole Audio Training Dataset Dataset description Luciole Audio Training Dataset is a large, multilingual, multi-task collection of audio–text conversations used to train the OpenLLM-France Luciole audio-language models. It adapts a text LLM to understand audio by pairing speech, music and environmental sounds with instruction-style dialogues (transcription, translation, spoken question answering, audio/music/sound captioning and question answering, speaker and… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Training-Dataset.textautomatic-speech-recognition10M<n<100M1 likes6 downloads19d agoHugging Face14OpenLLM-France /Luciole-Audio-Evaluation-Dataset Luciole Audio Evaluation Dataset Dataset description Luciole Audio Evaluation Dataset is the held-out evaluation collection used to benchmark the OpenLLM-France Luciole audio-language models. It is the evaluation counterpart of the Luciole Audio Training Dataset: a multilingual, multi-task set of audio–text conversations, restricted to the test splits of each source dataset, covering transcription, speech translation, spoken and music question answering, music… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/Luciole-Audio-Evaluation-Dataset.textautomatic-speech-recognition10K<n<100K1 likes6 downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.