Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shraavb /spanish-slang-stt-data Spanish Regional Speech-to-Text Dataset A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models. Dataset Description This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions: Region Samples Description Mexico 17,725 Mexican Spanish including CIEMPIESS corpus Spain 11,360 Castilian Spanish from TEDx and Common Voice Argentina 5,839 Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.audioautomatic-speech-recognition10K<n<100K0 likes4.4k downloads9mo agoHugging Face02skylar-ai-hf /skylar-dataset Skylar Dataset A curated speech dataset built for automatic speech recognition (ASR) benchmarking. It is assembled by streaming samples from existing public audio datasets and keeping only the ones that pass a fixed set of quality and diversity rules, organized by target language and audio duration. This is not a single-source dataset: samples are pulled from multiple upstream datasets into one unified structure. Re-running the pipeline with a different source (same or different… See the full description on the dataset page: https://huggingface.co/datasets/skylar-ai-hf/skylar-dataset.audioautomatic-speech-recognition1K<n<10K0 likes244 downloads10d agoHugging Face03bluesky7 /TEMA-Data TEMA-Data Data for TEMA: Evidence-Grounded Temporal Question Answering in Multi-Turn Multi-Audio Dialogs. Paper · Code · Models Datasets Component Config Split Size Temporal initialization temporal_init train / validation 98,401 / 500 examples TEMA-Dialog sft train 40,704 dialogues / 198,195 turns RL training rl_schedule train Download TEMA-Bench benchmark test 253 dialogues / 1,239 questions Tasks The 18-subtask guide maps the… See the full description on the dataset page: https://huggingface.co/datasets/bluesky7/TEMA-Data.texttext-generation10K<n<100K0 likes198 downloads16d agoHugging Face04rajjanardhan00 /Seamless_Dummy_Dataset_Fixed_3 MMLU-Pro json This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details. audioquestion-answeringn<1K0 likes163 downloads1y agoHugging Face05rorosese /my-voxtral-datasetaudion<1K0 likes158 downloads1y agoHugging Face06Codyfederer /tr-full-dataset TR-Full_dataset This is a merged speech dataset containing 41427 audio segments from 88 source datasets. Dataset Information Total Segments: 41427 Speakers: 222 Languages: tr Emotions: neutral, angry, sad, happy Original Datasets: 88 Dataset Structure Each example contains: audio: Audio file (WAV format, original sampling rate preserved) text: Transcription of the audio speaker_id: Unique speaker identifier (made unique across all merged… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/tr-full-dataset.audioautomatic-speech-recognition10K<n<100K6 likes112 downloads1y agoHugging Face07rodenhhh /ContextTTS_dataset ContextTTS Evaluation Dataset This is the official evaluation dataset for the paper "[ContextTTS Eval: A Benchmark for Evaluating Long-Form Contextual Expressive Text-to-Speech]". It is designed to evaluate the performance of multi-modal speech synthesis, specifically focusing on context-aware prosody and timbre consistency in Chinese conversations and audiobooks. Dataset Summary The dataset consists of high-quality Chinese audio-text pairs, organized into three distinct… See the full description on the dataset page: https://huggingface.co/datasets/rodenhhh/ContextTTS_dataset.audiotext-to-speechn<1K0 likes112 downloads6mo agoHugging Face08David-A-Amoo /naijavoices_dataset_85_hours_tts_bestFull&nbsp;dataset&nbsp;re-upload&nbsp;with&nbsp;more&nbsp;statistics&nbsp;as&nbsp;well&nbsp;as&nbsp;filtering&nbsp;scripts&nbsp;that&nbsp;give&nbsp;the&nbsp;top&nbsp;x&nbsp;files&nbsp;or&nbsp;best&nbsp;x&nbsp;hours&nbsp;for&nbsp;tts&nbsp;based&nbsp;and&nbsp;calculations&nbsp;from&nbsp;acoustic&nbsp;metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.tabular10K<n<100K2 likes89 downloads3mo agoHugging Face09Wi-Fi /korean-full-duplex-synthetic-dataset-preview Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.audioautomatic-speech-recognitionn<1K1 likes88 downloads2mo agoHugging Face10Aarjanm /youtube_audio_processed_dataset YouTube Audio Processed Dataset This dataset contains high-quality segmented speech datasets preprocessed from YouTube videos using the Emilia preprocessor framework. Dataset Structure Each row in the dataset contains a segmented audio clip, its aligned high-fidelity transcript, speaker labeling, and objective speech quality assessment scores. Features file_name: Audio column containing the relative path to the segmented .mp3 clip. text: The… See the full description on the dataset page: https://huggingface.co/datasets/Aarjanm/youtube_audio_processed_dataset.audion<1K0 likes88 downloads7d agoHugging Face11chtugha /small-german-medical-dialogue-dataset-for-moshi Small german dialogue dataset This dataset contains 500 completely made up medical phonecall dialogues between patients and a GP's office. Dataset Details Dataset Description 500 made up phonecalls that were first created with AI as text. The audio was then created using Openai tts-1-hd and the accurately timestamped transcripts were added. The audio files are formatted like this: Stereo with split channels: Speaker A is on the left channel… See the full description on the dataset page: https://huggingface.co/datasets/chtugha/small-german-medical-dialogue-dataset-for-moshi.audioaudio-text-to-textn<1K0 likes84 downloads4mo agoHugging Face12demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes78 downloads5mo agoHugging Face13ZHANGYUXUAN-zR /MCF-Dataset MCF: Text LLMS For Multimodal Emotional Causality Data Dataset task definition and annotation example of the MCF framework. The framework contains two core subtasks: five-tuple element extraction (identifying Target, Holder, Aspect, Opinion, Sentiment, and Rationale) and sentiment chain analysis (constructing causal relationship chains between emotional events). The dataset is provided with the following structure. Each sample includes video, audio, and dialogue… See the full description on the dataset page: https://huggingface.co/datasets/ZHANGYUXUAN-zR/MCF-Dataset.audiotext-classification10K<n<100K2 likes74 downloads1y agoHugging Face14TheMindExpansionNetwork /decision-lab-datasets-v0 Decision Lab — Synthetic GUI Decision Datasets (v0 sample batch) Synthetic training data for fine-tuning LiquidAI d1 decision models (d1-3B vision, d1-omni-600M audio) on GUI-screen and DJ-audio decision tasks. Generated 100% procedurally (no real user data, no scraped content, no AI image generation — just drawn rectangles, text, and synthesized tones). Part of: Sonic-Forage/decision-lab (private repo — code, configs, docs) What's inside vision/ —… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/decision-lab-datasets-v0.audio1K<n<10K0 likes73 downloads3d agoHugging Face15sunbv56 /song_dataset 🎵 Vietnamese Song Lyrics and Word Timestamps Dataset Dataset Summary The song_dataset provides high-quality Vietnamese song data, including metadata, full lyrics, and particularly word-level timestamps. This dataset is optimally designed for tasks such as: Training and evaluating automatic speech recognition (ASR) models on music. Lyrics synchronization (Lyrics Alignment / Karaoke generation). Natural language processing (NLP) analysis on song lyrics. The data is… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset.text1K<n<10K0 likes46 downloads7mo agoHugging Face16TumeloKonaite /synthetic-patient-dr-data Synthetic Patient DR Data Synthetic doctor-patient consultation dataset with structured clinical outputs and optional full-consultation audio. Dataset Summary This dataset was generated for research and prototyping in: clinical dialogue generation structured clinical extraction text-to-audio workflows conversational healthcare modeling All consultations are synthetic and should not be treated as real clinical encounters. Export Metadata Mode: audio Repo… See the full description on the dataset page: https://huggingface.co/datasets/TumeloKonaite/synthetic-patient-dr-data.audiotext-generationn<1K0 likes41 downloads6mo agoHugging Face17mira-iitjmu /ns-urdu-datasetaudio1K<n<10K0 likes30 downloads5mo agoHugging Face18anian0707 /scasr_datasetaudio10K<n<100K0 likes28 downloads2mo agoHugging Face19sunbv56 /song_dataset_chunked Vietnamese Songs Word-Level Timestamp Dataset (Chunked) This dataset contains word-level timestamp information for Vietnamese songs, specifically pre-chunked into segments up to 30 seconds for use in training or fine-tuning speech recognition (ASR) systems like Whisper. Dataset Summary The song_dataset_chunked provides high-quality Vietnamese song data, properly segmented into optimal ~30-second sequences. Duration Insights: Train split (train_chunked.jsonl): ~ 230.62… See the full description on the dataset page: https://huggingface.co/datasets/sunbv56/song_dataset_chunked.tabular10K<n<100K0 likes26 downloads7mo agoHugging Face20jeju-potato /jeju_potato_datasetsaudio10K<n<100K0 likes24 downloads1y agoHugging Face21RidheshBhati /Indic_New_dataset_TTS Indic TTS Dataset Hub (Mozilla) Validated audio–text pairs for multiple Indic languages from Mozilla Common Voice. Select the language from the Subset dropdown in the Dataset Viewer. Columns audio: WAV audio clip (16kHz) text: transcription duration: length in seconds speaking_rate: characters per second audio10K<n<100K0 likes21 downloads8mo agoHugging Face22Letian2003 /stage1a_smoke_data stage1a_smoke_data — AuT-ready 128-mel TFRecords (en/zh) Smoke-scale training data for Stage 1A input audio alignment of a Qwen3-ASR-AuT → MLP → frozen-VL-LLM omni model. Audio is pre-extracted 128-bin log-mel (the Qwen3-ASR AuT frontend: WhisperFeatureExtractor, 16 kHz, hop 160, n_fft 400) so training only needs to run the frozen AuT encoder — no raw-audio decoding at train time. 113,396 samples across 4 sources, stored as GZIP-compressed TFRecords (one file per source shard).… See the full description on the dataset page: https://huggingface.co/datasets/Letian2003/stage1a_smoke_data.tabularautomatic-speech-recognitionn<1K0 likes21 downloads3mo agoHugging Face23galammadin-asr /chlid-datasetaudio10K<n<100K0 likes19 downloads8mo agoHugging Face24electron-rare /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes17 downloads5mo agoHugging Face25IbraahimLab /voice-dataset Voice Dataset Collected from the web uploader tool. Voice Dataset Collected from the web uploader tool. audion<1K0 likes15 downloads8mo agoHugging Face26Ailiance-fr /mascarade-dsp-dataset Mascarade — DSP & Signal Processing Q&A ✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11) Per-sample Stack Exchange Electronics attribution recovered via the SE /search/advanced + /questions/{id} API search : 169 samples (~5.35 %) confirmed as Stack Exchange Electronics (CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution (URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60). 535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.texttext-generation1K<n<10K0 likes15 downloads5mo agoHugging Face27OpenDCAI /dataflow-mm-audio_asr_pipelineaudio1K<n<10K0 likes13 downloads8mo agoHugging Face28yadorigi /Onomatopoeia_Dataset🎧 Onomatopoeia Dataset (Audio → Manga Expression) 音声解析結果をもとに、日本語のオノマトペ(擬音語・擬態語)を生成するためのデータセットです。 本データセットは、音そのものではなく、音から推定された特徴・空間・情景を入力とする構造化データであり、 漫画的な表現生成を目的としたマルチモーダルデータです。 📌 Dataset Summary 本データセットは以下のパイプラインから生成されています: Audio ↓ Audio Features (04_features.json) ↓ Audio Events (05_audio_events.json) ↓ Space Judgement (06_space_judgement.json) ↓ Scene Interpretation (07_scene_interpretation.json) ↓ Onomatopoeia (08_onomatopoeia.json) 👉 音 → 空間 → 情景 → オノマトペ という段階的生成構造を持ちます。 📊… See the full description on the dataset page: https://huggingface.co/datasets/yadorigi/Onomatopoeia_Dataset.texttext-generationn<1K0 likes13 downloads7mo agoHugging Face29Diomande /s2o-datasetaudio10K<n<100K0 likes11 downloads11mo agoHugging Face30vichetkao /khmer_speech_news_datasetgated Khmer Speech Dataset Processing This repository contains scripts and instructions for preparing a Khmer speech dataset for machine learning tasks, such as automatic speech recognition (ASR). It demonstrates how to process a collection of audio files and metadata, and save them as Parquet files for efficient use in your training pipelines—without needing torchcodec. All dataset audio and transcripts in this project are sourced from https://wmc.org.kh/, the official website of… See the full description on the dataset page: https://huggingface.co/datasets/vichetkao/khmer_speech_news_dataset.audioaudio-classification10K<n<100K1 likes11 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.