Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01language-and-voice-lab /samromur_childrenThe Samrómur Children corpus contains more than 137000 validated speech-recordings uttered by Icelandic children.audioautomatic-speech-recognition10K<n<100K9 likes330 downloads3y agoHugging Face02language-and-voice-lab /samromur_milljonSamrómur Milljón consists of approximately 1 million of speech recordings (967 hours) collected through the platform samromur.is; the transcripts accompanying these recordings were automatically verified using various ASR systems such as: Wav2Vec, Whisper and NeMo.automatic-speech-recognition1M<n<10M4 likes261 downloads1y agoHugging Face03Sprakbanken /nb_samtaleNB Samtale is a speech corpus made by the Language Bank at the National Library of Norway. The corpus contains orthographically transcribed speech from podcasts and recordings of live events at the National Library. The corpus is intended as an open source dataset for Automatic Speech Recognition (ASR) development, and is specifically aimed at improving ASR systems’ handle on conversational speech.automatic-speech-recognition3 likes240 downloads3y agoHugging Face04language-and-voice-lab /samromur_asrSamrómur Icelandic Speech 1.0.automatic-speech-recognition100K<n<1M0 likes174 downloads4y agoHugging Face05Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes155 downloads5mo agoHugging Face06Sam04 /unfolded-veil-v9_traaudioautomatic-speech-recognition1K<n<10K0 likes109 downloads11mo agoHugging Face07OpenVoiceOS /stt-sampler-v1 stt-sampler-v1 Licensing: clips inherit their source dataset's license — CC-BY-4.0 for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE clips (source_dataset column identifies each clip's origin). A small, balanced, representative multilingual ASR eval sampler for the OVOS Plugin Arena: 100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32, one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")). Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.audioautomatic-speech-recognition1K<n<10K0 likes98 downloads2mo agoHugging Face08snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes97 downloads1y agoHugging Face09samil24 /kyrgyz-asr Kyrgyz Whisper Dataset Curated from local .wav + .wrd pairs (Arabic script) and normalized to Kyrgyz Cyrillic for ASR training. audioautomatic-speech-recognition10K<n<100K1 likes93 downloads1y agoHugging Face10juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes83 downloads7mo agoHugging Face11Indospeech /sundanese-spontaneous-conversation-samplegated Sundanese Spontaneous Conversation — Free Sample Unscripted two-speaker Sundanese (su-ID) conversation recorded in West Java, Indonesia. This is a free 30-minute sample of a larger commercially licensable corpus. Why this exists Sundanese has approximately 40 million native speakers, yet spontaneous conversational data is almost absent: Resource Type Volume Common Voice Spontaneous 4.0 Spontaneous, 2 speakers 0.62 h OpenSLR SLR36 / SLR44 Read speech —… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.audioautomatic-speech-recognitionn<1K1 likes79 downloads27d agoHugging Face12demegire /personaplex-finetuning-pharma-data-sample PersonaPlex Finetuning — Pharma Data Sample A 10-example slice of the synthetic patient-support / medication adherence dataset used to train demegire/personaplex-finetune-pharma. The on-disk layout below is exactly what the trainer in emotion-machine-org/personaplex-finetune consumes — use this as a template when building your own. Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at sample scale). Layout . ├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.audiotext-to-speechn<1K0 likes78 downloads5mo agoHugging Face13OcularAIInc /AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLEgated AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE Overview This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps. The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.audioautomatic-speech-recognitionn<1K0 likes71 downloads19d agoHugging Face14vikkyblacq /kare-codeswitch-samples Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.audioautomatic-speech-recognitionn<1K0 likes69 downloads25d agoHugging Face15Sin2pi /JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample. -Unnecessary and inaccurate punctuation have been removed. -Text has been normalized. Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja. audioautomatic-speech-recognition100K<n<1M9 likes54 downloads9mo agoHugging Face16jml2026 /reviewed_sample Silencio Network: Multilingual Accent Speech Dataset (Sample) Overview Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.audioautomatic-speech-recognition1K<n<10K0 likes53 downloads9mo agoHugging Face17snorbyte /indic-audio-natural-conversations-samplegated Dataset Card for Indic Audio Natural Conversations Sample Dataset Dataset Details Dataset Description The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.audioaudio-to-audion<1K1 likes49 downloads1y agoHugging Face18samuli /eduskunta-asr Finnish Parliament ASR Dataset Speech dataset from Finnish parliament (Eduskunta) plenary sessions with aligned ASR transcriptions and human reference texts. Overview Train Test Segments 64,945 9,121 Duration 423 h 59 h Speakers 212 187 Total: 74,066 segments, 482 hours, 212 unique speakers. Audio is 16 kHz mono Opus. Segments are individual speaker turns. Segments with CER > 0.30 are excluded (see below). Schema Column Type… See the full description on the dataset page: https://huggingface.co/datasets/samuli/eduskunta-asr.audioautomatic-speech-recognition10K<n<100K0 likes44 downloads6mo agoHugging Face19NicheVault /nichevault-afrikaans-asr-sample NicheVault Afrikaans ASR — Free Sample Overview NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations. This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.audioautomatic-speech-recognitionn<1K0 likes44 downloads3mo agoHugging Face20scubavoice /ibibio-efik-speech-corpus-sample Scuba Voice Dataset: Ibibio & Efik Sample (1 hour) Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria. This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.audioautomatic-speech-recognitionn<1K1 likes43 downloads3mo agoHugging Face21bhaskarvilles /clinical-speech-samples 🏥 Clinical Speech Samples Dataset High-Quality Clinical Audio Dataset for Medical AI Research Curated dataset for training and evaluating clinical speech processing models 📖 Dataset Description This dataset contains de-identified clinical audio samples for research and development of medical speech processing systems. It's designed to support training and evaluation of: 🎤 Speech enhancement in clinical settings 🔊 Audio source separation for medical… See the full description on the dataset page: https://huggingface.co/datasets/bhaskarvilles/clinical-speech-samples.automatic-speech-recognition1K<n<10K0 likes43 downloads3mo agoHugging Face22Luel-ai /luel-multilingual-tts-samplesgated Multilingual TTS Samples (Luel) License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE. A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.audiotext-to-speechn<1K0 likes42 downloads5mo agoHugging Face23Sam04 /dustlink-83p3_tran Sam04/dustlink-83p3_tran This dataset contains transcribed audio files organized in folders for scalability. Dataset Structure The dataset is organized with: Audio files: Stored in audio_XXXXX/ folders (5000 files per folder) Metadata: Stored in data_XXXXX/ folders as parquet files This organization follows Hugging Face best practices for datasets with millions of files. Statistics Total files: 1,199 Total batches: 1734 Audio folders: 2 Files per folder: max… See the full description on the dataset page: https://huggingface.co/datasets/Sam04/dustlink-83p3_tran.audioautomatic-speech-recognition10K<n<100K0 likes41 downloads11mo agoHugging Face24language-and-voice-lab /samromur_syntheticSamrómur Synthetic consists of 72 hours of synthetized speech in Icelandic.audioautomatic-speech-recognition10K<n<100K1 likes40 downloads2y agoHugging Face25MarieDeVox /saas-english-corporate-voice-dataset-sample PROFESSIONAL CONVERSATIONAL AI VOICE DATASET - SAAS CORPORATE SERIES Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 44.1kHz / 48kHz Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, human voice data optimized specifically for low-latency conversational AI UI/UX interfaces, intent-mapped software pipelines, and automated conversational SaaS call bots. This dataset is risk-free… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/saas-english-corporate-voice-dataset-sample.audioautomatic-speech-recognitionn<1K1 likes40 downloads4mo agoHugging Face26DexterityLearn /dexterity-pidgin-voice-dataset-sample-v2Dexterity Pidgin Voice Dataset — Sample v2 Creator: Dexterity Learn. Dataset Description 49 human-created Nigerian Pidgin English (Naija) utterances with aligned home studio audio, covering 8 domains: faith, wisdom, money, business, relationships, community, social commentary, and everyday conversation. Three utterances (U20–U22) address consent and anti-harassment themes, included for social good NLP use. What sets this dataset apart is its dual-translation structure — every utterance… See the full description on the dataset page: https://huggingface.co/datasets/DexterityLearn/dexterity-pidgin-voice-dataset-sample-v2.audioautomatic-speech-recognitionn<1K0 likes36 downloads4mo agoHugging Face27badrex /kinyarwanda-speech-sample Kinyarwanda Automatic Speech Recognition Dataset Dataset Description This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition. Dataset Details Language: Kinyarwanda (rw) Task: Automatic Speech Recognition Size: ~500 hours of transcribed speech Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads1y agoHugging Face28MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-sample-v1 BhasaFlow Khasi-English Parallel Sample v1 A professionally curated, gold-standard parallel speech and text corpus for the Khasi language. Published by Medharvix Systems Private Limited Part of the BhasaFlow Low-Resource Language Technology Initiative Overview This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.audioautomatic-speech-recognitionn<1K23 likes35 downloads6mo agoHugging Face29martinturuta /safi-diction-sample Safi Diction Sample This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents. The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours. This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.audioautomatic-speech-recognitionn<1K0 likes35 downloads1mo agoHugging Face30DavidErikMollberg /samromur_asr Dataset Card for samromur_asr Dataset Summary This is a modfied copy of the dataset from The Language and Voice Laboratory in RU. This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances. The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology. Languages The audio is in Icelandic. The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.audioautomatic-speech-recognition100K<n<1M0 likes34 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.