Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RodolfoRizzi /MyMentorLLM-dataset Dataset Card for MyMentorLLM This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information). Dataset Summary MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.audiotext-generation10K<n<100K0 likes1.3k downloads19d agoHugging Face02humair025 /hashed_data Munch Hashed Index - Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.tabulartext-generation1M<n<10M0 likes509 downloads10mo agoHugging Face03OmniEvalKit /omnievalkit-dataset OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 65 Total samples: 315,264 Total size: 620.3 GB (Parquet with embedded audio/image/video) Subsets with embedded video: 15 Subsets requiring external video download: 2 Usage from datasets import load_dataset ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.audioaudio-classification100K<n<1M0 likes413 downloads7mo agoHugging Face04IMJONEZZ /star-wars-dataset Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with the pipeline code (export_asr.py, export_diarization.py). The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.tabularautomatic-speech-recognition10K<n<100K3 likes345 downloads19d agoHugging Face05FBK-MT /fama-data Dataset Description, Collection, and Source The FAMA training data is the collection of English and Italian datasets for automatic speech recognition (ASR) and speech translation (ST) used to train the FAMA models family. The ASR section of FAMA is derived from the MOSEL data collection, including the automatic transcripts obtained with Whisper and available in the HuggingFace MOSEL Dataset. The ASR is further augmented with automatically transcribed speech from the… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/fama-data.tabulartranslation1M<n<10M2 likes315 downloads1y agoHugging Face06skylar-ai-hf /skylar-dataset Skylar Dataset A curated speech dataset built for automatic speech recognition (ASR) benchmarking. It is assembled by streaming samples from existing public audio datasets and keeping only the ones that pass a fixed set of quality and diversity rules, organized by target language and audio duration. This is not a single-source dataset: samples are pulled from multiple upstream datasets into one unified structure. Re-running the pipeline with a different source (same or different… See the full description on the dataset page: https://huggingface.co/datasets/skylar-ai-hf/skylar-dataset.audioautomatic-speech-recognition1K<n<10K0 likes244 downloads10d agoHugging Face07Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes170 downloads1mo agoHugging Face08paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes161 downloads17d agoHugging Face09martinnavs /ai231-fil-supplemental-data AI231 ME2 Filipino-accented synthetic commands (supplemental training data) Optional extra training data for the AI231 ME2 voice-command dataset: English voice commands spoken by 17 cloned Filipino-accented voices. It follows that dataset's schema and is train only: it is not part of its train, test or holdout splits. data/train-0000{0..3}-of-00004.parquet 14,120 clips, 8.7 h, 17 voices, 19 commands (223-1,482 clips each) Audio: 16 kHz, mono, 16-bit PCM WAV, stored in the… See the full description on the dataset page: https://huggingface.co/datasets/martinnavs/ai231-fil-supplemental-data.audioautomatic-speech-recognition10K<n<100K0 likes94 downloads7d agoHugging Face10Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes88 downloads2mo agoHugging Face11FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes60 downloads5mo agoHugging Face12kvest /Swedia-ASR-Dataset Swedia ASR Dataset This repository contains a small Swedish ASR evaluation dataset based on speech transcriptions from Swedia 2000. It was assembled to compare automatic speech-recognition output against manually corrected reference transcriptions for Swedish dialectal speech. The dataset is useful for quick experiments with Swedish ASR systems, especially when you want to inspect recognition quality on spontaneous speech from different regions, speakers, ages, and genders.… See the full description on the dataset page: https://huggingface.co/datasets/kvest/Swedia-ASR-Dataset.tabularautomatic-speech-recognitionn<1K1 likes46 downloads5mo agoHugging Face13anuj-inavlabs /kupe-asr-en-data kupe-asr-en-mini-150m — data Two loadable configs, packed into ~20-25 bunch_*.parquet files each (Hub-quota friendly): raw — 24 kHz mono English audio (flac bytes) + text. The encode stage reads this. mimi — Mimi c0..c7 codes (12.5 Hz) + text. Training reads this. Ledgers under ledger/ (data.json, mimi.json) track collected/encoded hours and resume state. from datasets import load_dataset ds = load_dataset("anuj-inavlabs/kupe-asr-en-data", "mimi", split="train") tabularautomatic-speech-recognition1M<n<10M0 likes39 downloads1mo agoHugging Face14ayush712145 /sarvam-tts-dataset Sarvam TTS Dataset: Indian English and Hindi A curated speech dataset for TTS model training, containing 396 clips in Indian English (en-IN) and Hindi (hi-IN) totalling 57.3 minutes (3440 seconds). Pipeline source code: https://github.com/Ayush147258/sarvam-tts-dataset Dataset Summary Stat Value Total clips 396 Total duration 57.3 minutes (3440 seconds) Clip duration range 5.0s to 26.5s Mean clip duration 8.7s Median clip duration 7.5s… See the full description on the dataset page: https://huggingface.co/datasets/ayush712145/sarvam-tts-dataset.tabulartext-to-speechn<1K0 likes29 downloads4mo agoHugging Face15DrUkachi /ktt-math-tutor-data KTT Math Tutor — Data Data artefacts for the AIMS KTT Hackathon Tier-3 submission S2.T3.1 AI Math Tutor for Early Learners. Source code: https://github.com/DrUkachi/ktt-math-tutor. Contents T3.1_Math_Tutor/ Core curriculum + seeds. curriculum.json — 80 items × 5 sub-skills (counting, number sense, addition, subtraction, word problem) with EN / FR / KIN stems, difficulty 1–10, age bands 5–6 / 6–7 / 7–8 / 8–9, visual asset keys, expected integer answer.… See the full description on the dataset page: https://huggingface.co/datasets/DrUkachi/ktt-math-tutor-data.tabularquestion-answeringn<1K0 likes28 downloads6mo agoHugging Face16PalakEngineerMaster /Processed_TTS_Multilingual_Data Processed TTS Multilingual Data Validated and quality-checked multilingual speech datasets for TTS training, covering 12+ Indian languages. Datasets Included Subset Samples Hours Description indic_voices_r 239,684 548.8h Indic Voices_R — IVR recordings rasa 201,509 361.2h RASA — read speech (wiki, conv, book, news) indictts_iitm 155,236 253.6h Indic TTS (IIT Madras) — studio TTS recordings at 48kHz Total 596,429 1,163.6h Languages… See the full description on the dataset page: https://huggingface.co/datasets/PalakEngineerMaster/Processed_TTS_Multilingual_Data.tabulartext-to-speech100K<n<1M0 likes25 downloads8mo agoHugging Face17Sakchham19 /Trump_Voice_Dataset Trump Voice Dataset This dataset contains audio clips of Donald Trump's speech from the World Economic Forum (WEF) 2018, paired with their corresponding transcriptions. The dataset is designed for text-to-speech (TTS) and speech recognition tasks. Dataset Description Dataset Summary The Trump Voice Dataset consists of 20 audio samples (10 train, 10 test) extracted from Donald Trump's speech at the World Economic Forum 2018. Each audio clip is approximately 10… See the full description on the dataset page: https://huggingface.co/datasets/Sakchham19/Trump_Voice_Dataset.tabularautomatic-speech-recognitionn<1K0 likes23 downloads1y agoHugging Face18Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes21 downloads3mo agoHugging Face19Khaledtelbahnasy /egyptian-arabic-stt-data Egyptian Arabic STT Dataset Synthetic Egyptian Arabic speech dataset generated by the Synthetic Egyptian Speech Data Pipeline. Samples are human-reviewed and quality-validated using Whisper ASR (WER/CER). Dataset Statistics Metric Value Total samples 50 Total duration 85.2s (0.02h) Dialect validated 50 / 50 Average WER 0.4136 Average CER 0.1642 Topics food_ordering Fields Field Type Description id string Deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Khaledtelbahnasy/egyptian-arabic-stt-data.tabularautomatic-speech-recognitionn<1K1 likes20 downloads5mo agoHugging Face20Pedramebd /welsh-speech-dataset Welsh Speech Dataset A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations. Dataset Overview Speakers: 33 participants Phrases: 10 Welsh phrases per speaker Sequences: ~330 (33 speakers x 10 phrases) Modalities: Audio recordings (.wav) 3D facial reconstructions (.obj meshes + texture maps) 68-point facial landmarks (ibug68 template) Fluency Scores: Each phrase rated 0-5… See the full description on the dataset page: https://huggingface.co/datasets/Pedramebd/welsh-speech-dataset.tabularautomatic-speech-recognitionn<1K0 likes19 downloads2mo agoHugging Face21arvinsingh /welsh-speech-dataset Welsh Speech Dataset A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations. Dataset Overview Speakers: 33 participants Phrases: 10 Welsh phrases per speaker Sequences: ~330 (33 speakers x 10 phrases) Modalities: Audio recordings (.wav) 3D facial reconstructions (.obj meshes + texture maps) 68-point facial landmarks (ibug68 template) Fluency Scores: Each phrase rated 0-5 (5 =… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-dataset.tabularautomatic-speech-recognitionn<1K0 likes13 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.