Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenVoiceOS /stt-sampler-v1 stt-sampler-v1 Licensing: clips inherit their source dataset's license — CC-BY-4.0 for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE clips (source_dataset column identifies each clip's origin). A small, balanced, representative multilingual ASR eval sampler for the OVOS Plugin Arena: 100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32, one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")). Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.audioautomatic-speech-recognition1K<n<10K0 likes98 downloads2mo agoHugging Face02snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes97 downloads1y agoHugging Face03juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes83 downloads7mo agoHugging Face04Indospeech /sundanese-spontaneous-conversation-samplegated Sundanese Spontaneous Conversation — Free Sample Unscripted two-speaker Sundanese (su-ID) conversation recorded in West Java, Indonesia. This is a free 30-minute sample of a larger commercially licensable corpus. Why this exists Sundanese has approximately 40 million native speakers, yet spontaneous conversational data is almost absent: Resource Type Volume Common Voice Spontaneous 4.0 Spontaneous, 2 speakers 0.62 h OpenSLR SLR36 / SLR44 Read speech —… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.audioautomatic-speech-recognitionn<1K1 likes79 downloads27d agoHugging Face05demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes78 downloads5mo agoHugging Face06OcularAIInc /AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLEgated AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE Overview This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps. The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.audioautomatic-speech-recognitionn<1K0 likes71 downloads18d agoHugging Face07vikkyblacq /kare-codeswitch-samples Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.audioautomatic-speech-recognitionn<1K0 likes69 downloads25d agoHugging Face08Sin2pi /JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample. -Unnecessary and inaccurate punctuation have been removed. -Text has been normalized. Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja. audioautomatic-speech-recognition100K<n<1M9 likes54 downloads9mo agoHugging Face09jml2026 /reviewed_sample Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.audioautomatic-speech-recognition1K<n<10K0 likes53 downloads9mo agoHugging Face10snorbyte /indic-audio-natural-conversations-samplegated Dataset Card for Indic Audio Natural Conversations Sample Dataset Dataset Details Dataset Description The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.audioaudio-to-audion<1K1 likes49 downloads1y agoHugging Face11NicheVault /nichevault-afrikaans-asr-sample NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.audioautomatic-speech-recognitionn<1K0 likes44 downloads3mo agoHugging Face12scubavoice /ibibio-efik-speech-corpus-sample Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.audioautomatic-speech-recognitionn<1K1 likes43 downloads3mo agoHugging Face13bhaskarvilles /clinical-speech-samples 🏥 Clinical Speech Samples Dataset High-Quality Clinical Audio Dataset for Medical AI Research Curated dataset for training and evaluating clinical speech processing models 📖 Dataset Description This dataset contains de-identified clinical audio samples for research and development of medical speech processing systems. It's designed to support training and evaluation of: 🎤 Speech enhancement in clinical settings 🔊 Audio source separation for medical… See the full description on the dataset page: https://huggingface.co/datasets/bhaskarvilles/clinical-speech-samples.automatic-speech-recognition1K<n<10K0 likes43 downloads3mo agoHugging Face14Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes42 downloads5mo agoHugging Face15MarieDeVox /saas-english-corporate-voice-dataset-sample PROFESSIONAL CONVERSATIONAL AI VOICE DATASET - SAAS CORPORATE SERIES Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 44.1kHz / 48kHz Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, human voice data optimized specifically for low-latency conversational AI UI/UX interfaces, intent-mapped software pipelines, and automated conversational SaaS call bots. This dataset is risk-free… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/saas-english-corporate-voice-dataset-sample.audioautomatic-speech-recognitionn<1K1 likes40 downloads4mo agoHugging Face16DexterityLearn /dexterity-pidgin-voice-dataset-sample-v2Dexterity Pidgin Voice Dataset — Sample v2 Creator: Dexterity Learn. Dataset Description 49 human-created Nigerian Pidgin English (Naija) utterances with aligned home studio audio, covering 8 domains: faith, wisdom, money, business, relationships, community, social commentary, and everyday conversation. Three utterances (U20–U22) address consent and anti-harassment themes, included for social good NLP use. What sets this dataset apart is its dual-translation structure — every utterance… See the full description on the dataset page: https://huggingface.co/datasets/DexterityLearn/dexterity-pidgin-voice-dataset-sample-v2.audioautomatic-speech-recognitionn<1K0 likes36 downloads4mo agoHugging Face17badrex /kinyarwanda-speech-sample Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads1y agoHugging Face18MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K23 likes35 downloads6mo agoHugging Face19martinturuta /safi-diction-sample Safi Diction Sample This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents. The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours. This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.audioautomatic-speech-recognitionn<1K0 likes35 downloads1mo agoHugging Face20Veronica1NW /cv17_sw_kenyan_sample Common Voice 17.0 — Swahili (Kenyan Sample) This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices. It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili. Dataset Summary Language: Kiswahili (Swahili, sw) Accent/Region: Kenyan speakers Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.audioautomatic-speech-recognition1K<n<10K0 likes34 downloads1y agoHugging Face21Cybrpgs /corpus5-t2inserts-sample-200-20260916gatedaudioautomatic-speech-recognitionn<1K0 likes31 downloads24d agoHugging Face22BrunoHays /muscat-merged-samples MUSCAT — Merged Long-Form Samples This dataset is a merged, long-form reformatting of goodpiku/muscat-eval (MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation Evaluation). The original MUSCAT release stores each conversation as many short, single-language segments. Here those segments are concatenated back into one continuous recording per conversation, so each row is a single long-form code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.audioautomatic-speech-recognitionn<1K0 likes30 downloads1mo agoHugging Face231infinity0 /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K1 likes28 downloads6mo agoHugging Face24psdn-ai /mandarin-speech-samplesgated Mandarin Speech Samples This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery. What This Shows Mandarin speech audio with consistent metadata Clip-level transcript fields for content review Language and format signals for procurement review Dataset Specifications Field Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.audioautomatic-speech-recognitionn<1K0 likes28 downloads3mo agoHugging Face25archivartaunik /be-sidon-restored-sample-100 be-sidon-restored-sample-100 Набор прыкладаў беларускай мовы з Common Voice (validated), апрацаваны мадэллю аднаўлення (denoise). Файлы audio/original/ — арыгінальныя кліпы (як у Common Voice). audio/restored/ — адноўленыя WAV (48 кГц). metadata.csv — палеткі: audio, original_audio, sentence, speaker. Дата стварэння: 2025-09-25. automatic-speech-recognitionn<1K0 likes26 downloads1y agoHugging Face26Cybrpgs /corpus5-t2inserts-sample-1000-20260916gatedaudioautomatic-speech-recognition1K<n<10K0 likes26 downloads24d agoHugging Face27snorbyte /indic-tts-sample-snac-encodedgated Dataset Card for Indic TTS Sample SNAC Encoded Dataset Dataset Details Dataset Description The IndicTTSSampleSNACEncoded Dataset is a multilingual, text-speech pair sample dataset. It features ~135 hours of human-voiced recordings of transcripts in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi, across multiple speakers and other metadata. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-tts-sample-snac-encoded.textaudio-to-audio10K<n<100K3 likes25 downloads1y agoHugging Face28BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes25 downloads5mo agoHugging Face29Cybrpgs /stillalive-sample-1000-uniform-20260919gated stillalive sample 1000 (uniform over 625,217) 1000 chunks drawn uniformly at random, without replacement, from all 625,217 verified chunks merged so far (batches saq_b0001..saq_b0040, partA final). Same schema as the full dataset (audio embedded in the audio struct column). Distinct recordings: 913; audio: 2.01 h Randomness: secrets.SystemRandom().sample (os.urandom CSPRNG, Fisher-Yates). No seed. Audit: sample_manifest.json holds sha256 over the sorted (recording_id, chunk_id)… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-uniform-20260919.audioautomatic-speech-recognition1K<n<10K0 likes25 downloads21d agoHugging Face30ODYSSEYAILABS /odyssey-sa-voice-v0.4-sampleOdyssey SA Voice Corpus (V0.4) — Evaluation Preview Overview The Odyssey SA Voice Corpus (V0.4) is a 1000-hour multilingual South African speech dataset spanning 8 languages, designed for automatic speech recognition (ASR), code-switching research, and large audio model (LAM) evaluation. This repository provides a 69-minute evaluation preview (Episode 001 — Nobantu Vilakazi). The preview mirrors the segmentation, metadata schema, validation pipeline, and structural standards used in the full… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-sa-voice-v0.4-sample.audioautomatic-speech-recognitionn<1K0 likes22 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.