Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HTH-inc /japanese-casual-conversational-speech-golden-dataset-preview Japanese Casual Conversational Speech Golden Dataset (Preview) 💼 Commercial License & Full Access This repository contains a limited preview. The full 60-hour dataset collected via the "Kataro" app is available for commercial use, ASR benchmarking, and Spoken Dialogue Model fine-tuning. To purchase the full dataset, please contact us: 👉 Email: info@hth-inc.com 👉 Website: https://hth-inc.com/business 🌟 4 Reasons to Choose This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HTH-inc/japanese-casual-conversational-speech-golden-dataset-preview.audioautomatic-speech-recognitionn<1K2 likes417 downloads1mo agoHugging Face02baptistefrancois1 /conversational-s2s-v1-en Conversational speech-to-speech dialogues (re-voiced) Synthetic spoken multi-turn conversations between a user and a voice assistant, built to finetune a speech-to-speech model (LFM2-Audio). This set holds the dialogues of Rcarvalo/conversational-s2s-v1@6fc04ee, word for word: only the voices differ. Every spoken turn was synthesized again, so the clips of the source do not mix with these. A dialogue is a list of spoken turns (user, assistant, user, ...), each with its text and… See the full description on the dataset page: https://huggingface.co/datasets/baptistefrancois1/conversational-s2s-v1-en.audio-to-audio1K<n<10K0 likes305 downloads4d agoHugging Face03Makan09 /bam-asr-conversational All Bambara ASR Dataset This is the dataset that fueled our early ASR experiments that gave as results the V0 models. It is primarily composed of the Jeli-ASR dataset (available at RobotsMali/jeli-asr), along with the Mali-Pense data curated and published by Aboubacar Ouattara (available at oza75/bambara-tts). Additionally, it includes 1 hour of audio recently collected by the RobotsMali AI4D Lab, featuring children's voices reading some of RobotsMali GAIFE books. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/bam-asr-conversational.audioautomatic-speech-recognition10K<n<100K2 likes164 downloads1mo agoHugging Face04jml2026 /conversational-speech-dataset 🎙️ Silencio Network: Conversational Speech Dataset Overview Sample conversational speech data from Silencio Network's crowdsourced voice AI platform. This dataset contains multi-speaker meeting recordings with word-level transcripts, speaker diarization, and rich demographic metadata. Each row represents one participant in a meeting and includes 3 audio files: Audio Column Description Format file_name (speaker audio) Individual participant's… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/conversational-speech-dataset.automatic-speech-recognitionn<1K0 likes116 downloads7mo agoHugging Face05beatsprom /realtime-conversational-voice-agent-duplex-2026 🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026) This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab. The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.texttext-generationn<1K0 likes78 downloads1mo agoHugging Face06OcularAIInc /AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLEgated AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE Overview This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps. The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.audioautomatic-speech-recognitionn<1K0 likes71 downloads19d agoHugging Face07UsergyAI /Global-Conversational-Speechgated Global Conversational Speech Dataset 305 hours. 18 locales. Real conversations. Not scraped from YouTube. Not recorded by anonymous crowds who don't speak the language. Every conversation in this dataset traces back to verified native speakers we know by name. The [Human] Standard Most speech datasets are built the same way: scrape the internet, hire anonymous contractors, run it through automated QC, ship it. The result? Models that are confidently wrong. We… See the full description on the dataset page: https://huggingface.co/datasets/UsergyAI/Global-Conversational-Speech.audioautomatic-speech-recognition100K<n<1M3 likes65 downloads3mo agoHugging Face08AirCaps /mega-asr-conversational-overlap Mega-ASR Conversational Overlap Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set derived from AirCaps/mega-asr-noise-a5sv2, which in turn is sampled from the Mega-ASR training corpus zhifeixie/Voices-in-the-Wild-2M. The existing AirCaps dataset evaluates single-utterance acoustic robustness. This companion dataset evaluates a different failure mode: two-turn conversational continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.audioautomatic-speech-recognitionn<1K0 likes53 downloads2mo agoHugging Face09Squadstack /conversational-streaming-asr-benchmark SquadStack Conversational Streaming ASR Benchmark (8 kHz) Version 1.0.0 · maintained by SquadStack Schema · Leaderboard · Latency · Submit a system · Licence · Terms of use Key takeaways What this is. 863 real Hindi–English telesales calls (5.53 hours of customer speech, 8 kHz phone audio), human-transcribed turn by turn, and 11 speech recognisers scored on them. The question it answers: which recogniser should run inside an Indian voice agent, judged on… See the full description on the dataset page: https://huggingface.co/datasets/Squadstack/conversational-streaming-asr-benchmark.audioautomatic-speech-recognition1K<n<10K1 likes49 downloads1d agoHugging Face10Datoric /tts-conversational-voice-20000hgated TTS Voice Dataset 20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI. This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples. Overview… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/tts-conversational-voice-20000h.tabulartext-to-speech100K<n<1M0 likes48 downloads3mo agoHugging Face11DatoricAI /tts-conversational-voice-20000hgated TTS Voice Dataset 20,000 hours of high-fidelity 48kHz conversational audio across 30+ global, regional, and underrepresented languages, built for text-to-speech, voice cloning, and multilingual speech AI. This repository is a specification and preview listing. The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full schema and get real audio samples. Overview The TTS Voice Dataset is a 20,000-hour collection of… See the full description on the dataset page: https://huggingface.co/datasets/DatoricAI/tts-conversational-voice-20000h.text-to-speech1M<n<10M0 likes35 downloads3mo agoHugging Face12ShimogaAIteam /conversational_kannada_stt Conversational Kannada STT This dataset contains corrected transcriptions of conversational Kannada speech,prepared specifically for fine-tuning Whisper models on conversational and dialectal Kannada. Unlike many ASR datasets, this release provides pre-computed Whisper input features (log-Mel spectrograms)so you can train/fine-tune Whisper models without raw audio processing. Dataset Creation Source Audio: Publicly available YouTube videos in Kannada. Initial… See the full description on the dataset page: https://huggingface.co/datasets/ShimogaAIteam/conversational_kannada_stt.textautomatic-speech-recognition1K<n<10K1 likes15 downloads1y agoHugging Face13SpeechDataAI /egyptian-arabic-conversational-speech Egyptian Arabic Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in Egyptian Arabic, recorded in Egypt with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/egyptian-arabic-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face14SpeechDataAI /us-english-conversational-speech US English Conversational Speech Dataset 1,500 hours of spontaneous two-person conversations in US English, recorded in the United States with 150 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/us-english-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face15SpeechDataAI /south-african-english-conversational-speech South African English Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in South African English, recorded in South Africa with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/south-african-english-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face16SpeechDataAI /singapore-english-conversational-speech Singapore English Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in Singapore English, recorded in Singapore with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/singapore-english-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face17SpeechDataAI /european-spanish-conversational-speech European Spanish Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in European Spanish, recorded in Spain with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/european-spanish-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face18SpeechDataAI /latin-american-spanish-conversational-speech Latin American Spanish Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in Latin American Spanish, recorded in Mexico with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/latin-american-spanish-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face19SpeechDataAI /french-conversational-speech French Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in French, recorded in France with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent on… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/french-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face20SpeechDataAI /congolese-french-conversational-speech Congolese French Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in Congolese French, recorded in the DR Congo with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/congolese-french-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face21SpeechDataAI /european-portuguese-conversational-speech European Portuguese Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in European Portuguese, recorded in Portugal with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/european-portuguese-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face22SpeechDataAI /brazilian-portuguese-conversational-speech Brazilian Portuguese Conversational Speech Dataset 2,000 hours of spontaneous two-person conversations in Brazilian Portuguese, recorded in Brazil with 200 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/brazilian-portuguese-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face23SpeechDataAI /angolan-portuguese-conversational-speech Angolan Portuguese Conversational Speech Dataset 500 hours of spontaneous two-person conversations in Angolan Portuguese, recorded in Angola with 50 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/angolan-portuguese-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face24SpeechDataAI /german-conversational-speech German Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in German, recorded in Germany with 50 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent on… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/german-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face25SpeechDataAI /dutch-conversational-speech Dutch Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in Dutch, recorded in the Netherlands with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/dutch-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face26SpeechDataAI /surinamese-dutch-conversational-speech Surinamese Dutch Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in Surinamese Dutch, recorded in Suriname with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/surinamese-dutch-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face27SpeechDataAI /italian-conversational-speech Italian Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in Italian, recorded in Italy with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent on… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/italian-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face28SpeechDataAI /polish-conversational-speech Polish Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in Polish, recorded in Poland with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent on… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/polish-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face29SpeechDataAI /russian-conversational-speech Russian Conversational Speech Dataset 1,000 hours of spontaneous two-person conversations in Russian, recorded in Russia with 100 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent on… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/russian-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face30SpeechDataAI /ukrainian-conversational-speech Ukrainian Conversational Speech Dataset 1,500 hours of spontaneous two-person conversations in Ukrainian, recorded in Ukraine with 150 vetted native speakers. Dual-channel audio, time-aligned human-verified transcripts and speaker metadata, licensed for commercial ASR, TTS and voice AI training. This repository is the dataset card only. The audio is licensed commercially and isn't hosted on Hugging Face. Free samples (audio, matching transcripts and the metadata schema) are sent… See the full description on the dataset page: https://huggingface.co/datasets/SpeechDataAI/ukrainian-conversational-speech.audioautomatic-speech-recognition0 likes7h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.