datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stt-sampler-v1
stt-sampler-v1
Licensing: clips inherit their source dataset's license — CC-BY-4.0
for MInDS-14 and FLEURS clips, CC-BY-NC-SA-4.0 for Speech-MASSIVE
clips (source_dataset column identifies each clip's origin).
A small, balanced, representative multilingual ASR eval sampler for the
OVOS Plugin Arena:
100 clips per language x 20 locales = 2000 clips, 16 kHz mono float32,
one config per locale (load_dataset("OpenVoiceOS/stt-sampler-v1", "<lang>")).
Designed to seed every STT… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/stt-sampler-v1.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.sundanese-spontaneous-conversation-sample
Sundanese Spontaneous Conversation — Free Sample
Unscripted two-speaker Sundanese (su-ID) conversation recorded in
West Java, Indonesia. This is a free 30-minute sample of a larger
commercially licensable corpus.
Why this exists
Sundanese has approximately 40 million native speakers, yet
spontaneous conversational data is almost absent:
Resource
Type
Volume
Common Voice Spontaneous 4.0
Spontaneous, 2 speakers
0.62 h
OpenSLR SLR36 / SLR44
Read speech
—… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE
AMERICAN ENGLISH TRANSCRIBED HI-FI FULL-DUPLEX TWO-SPEAKER CONVERSATIONAL DATASET — SAMPLE
Overview
This open sample from Ocular AI contains four American English conversations between two people, with a separate audio track for each speaker and verbatim transcripts containing segment- and word-level timestamps.
The recordings capture conversational exchanges: repetitions, fillers, false starts, pauses, laughter, and audible breaths. Some conversations begin with… See the full description on the dataset page: https://huggingface.co/datasets/OcularAIInc/AMERICAN-ENGLISH-TRANSCRIBED-HIFI-FULL-DUPLEX-TWO-SPEAKER-CONVERSATIONAL-DATASET-SAMPLE.kare-codeswitch-samples
Kare — Code-Switching Illustrative Samples
Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin
code-switching that Kare, a voice-first
AI health assistant for Nigeria, is built to understand — submitted as part
of Kare's entry to the Sahara CodeSwitch Africa Challenge.
What this is — and isn't
Is: eight original sentences, written by the Kare team specifically
for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian
Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.indic-audio-natural-conversations-sample
Dataset Card for Indic Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.nichevault-afrikaans-asr-sample
NicheVault Afrikaans ASR — Free Sample
Overview
NicheVault Afrikaans ASR — Free Sample is a 20-clip preview of a 61.5-hour Whisper-ready Afrikaans ASR training dataset. Every clip is 16kHz mono WAV with a human-verified sentence-level transcript, formatted as JSONL metadata. All clips are pure CC BY — no ShareAlike, no copyleft obligations.
This sample is drawn from 20 clips across the full dataset's train/validation/test splits. It is provided unauthenticated… See the full description on the dataset page: https://huggingface.co/datasets/NicheVault/nichevault-afrikaans-asr-sample.ibibio-efik-speech-corpus-sample
Scuba Voice Dataset: Ibibio & Efik Sample (1 hour)
Scuba is a voice data infrastructure company collecting, labeling, and licensing speech corpora for low-resource languages. This repository contains a 1-hour sample of labeled Ibibio and Efik speech data collected from native speakers in Akwa Ibom, Nigeria.
This sample is provided for evaluation purposes only. For access to the full dataset (~200 hours across Ibibio and Efik), contact us at info@scubavoice.com.… See the full description on the dataset page: https://huggingface.co/datasets/scubavoice/ibibio-efik-speech-corpus-sample.clinical-speech-samples
🏥 Clinical Speech Samples Dataset
High-Quality Clinical Audio Dataset for Medical AI Research
Curated dataset for training and evaluating clinical speech processing models
📖 Dataset Description
This dataset contains de-identified clinical audio samples for research and development of medical speech processing systems. It's designed to support training and evaluation of:
🎤 Speech enhancement in clinical settings
🔊 Audio source separation for medical… See the full description on the dataset page: https://huggingface.co/datasets/bhaskarvilles/clinical-speech-samples.luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.saas-english-corporate-voice-dataset-sample
PROFESSIONAL CONVERSATIONAL AI VOICE DATASET - SAAS CORPORATE SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 44.1kHz / 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, human voice data optimized specifically for low-latency conversational AI UI/UX interfaces, intent-mapped software pipelines, and automated conversational SaaS call bots. This dataset is risk-free… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/saas-english-corporate-voice-dataset-sample.dexterity-pidgin-voice-dataset-sample-v2Dexterity Pidgin Voice Dataset — Sample v2
Creator: Dexterity Learn.
Dataset Description
49 human-created Nigerian Pidgin English (Naija) utterances with aligned home studio audio, covering 8 domains: faith, wisdom, money, business, relationships, community, social commentary, and everyday conversation. Three utterances (U20–U22) address consent and anti-harassment themes, included for social good NLP use.
What sets this dataset apart is its dual-translation structure — every utterance… See the full description on the dataset page: https://huggingface.co/datasets/DexterityLearn/dexterity-pidgin-voice-dataset-sample-v2.kinyarwanda-speech-sample
Kinyarwanda Automatic Speech Recognition Dataset
Dataset Description
This dataset contains a sample from the 500 hours of Kinyarwanda speech data covering Health, Government, Finance, Education, and Agriculture domains, converted from the Kaggle Kinyarwanda ASR Track A competition.
Dataset Details
Language: Kinyarwanda (rw)
Task: Automatic Speech Recognition
Size: ~500 hours of transcribed speech
Domains: Health, Government, Financial Services, Education… See the full description on the dataset page: https://huggingface.co/datasets/badrex/kinyarwanda-speech-sample.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-sample-v1.safi-diction-sample
Safi Diction Sample
This dataset is a sample speech dataset containing short audio recordings paired with expected transcription text from respondents.
The dataset was created to show a sample of the type of data that can be collected with Safi's collection engine. Responses were collected remotely within the span of 12 hours.
This is just a sample for development or testing - contact hq@safidata.com for the full dataset or custom datasets, or visit https://www.safidata.com/.… See the full description on the dataset page: https://huggingface.co/datasets/martinturuta/safi-diction-sample.cv17_sw_kenyan_sample
Common Voice 17.0 — Swahili (Kenyan Sample)
This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices.
It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili.
Dataset Summary
Language: Kiswahili (Swahili, sw)
Accent/Region: Kenyan speakers
Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.corpus5-t2inserts-sample-200-20260916muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.bhasaflow-khasi-english-parallel-sample-v1
BhasaFlow Khasi-English Parallel Sample v1
A professionally curated, gold-standard parallel speech and text corpus for the Khasi language.
Published by Medharvix Systems Private Limited
Part of the BhasaFlow Low-Resource Language Technology Initiative
Overview
This repository contains a public sample preview of the BhasaFlow Khasi-English Parallel Corpus, a structured speech and text dataset developed by Medharvix Systems Private Limited. The dataset pairs… See the full description on the dataset page: https://huggingface.co/datasets/1infinity0/bhasaflow-khasi-english-parallel-sample-v1.mandarin-speech-samples
Mandarin Speech Samples
This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery.
What This Shows
Mandarin speech audio with consistent metadata
Clip-level transcript fields for content review
Language and format signals for procurement review
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.be-sidon-restored-sample-100
be-sidon-restored-sample-100
Набор прыкладаў беларускай мовы з Common Voice (validated), апрацаваны мадэллю аднаўлення (denoise).
Файлы
audio/original/ — арыгінальныя кліпы (як у Common Voice).
audio/restored/ — адноўленыя WAV (48 кГц).
metadata.csv — палеткі: audio, original_audio, sentence, speaker.
Дата стварэння: 2025-09-25.
corpus5-t2inserts-sample-1000-20260916indic-tts-sample-snac-encoded
Dataset Card for Indic TTS Sample SNAC Encoded Dataset
Dataset Details
Dataset Description
The IndicTTSSampleSNACEncoded Dataset is a multilingual, text-speech pair sample dataset. It features ~135 hours of human-voiced recordings of transcripts in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi, across multiple speakers and other metadata.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-tts-sample-snac-encoded.english-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.stillalive-sample-1000-uniform-20260919
stillalive sample 1000 (uniform over 625,217)
1000 chunks drawn uniformly at random, without replacement, from all 625,217 verified
chunks merged so far (batches saq_b0001..saq_b0040, partA final). Same schema as the full
dataset (audio embedded in the audio struct column).
Distinct recordings: 913; audio: 2.01 h
Randomness: secrets.SystemRandom().sample (os.urandom CSPRNG, Fisher-Yates). No seed.
Audit: sample_manifest.json holds sha256 over the sorted (recording_id, chunk_id)… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/stillalive-sample-1000-uniform-20260919.odyssey-sa-voice-v0.4-sampleOdyssey SA Voice Corpus (V0.4) — Evaluation Preview
Overview
The Odyssey SA Voice Corpus (V0.4) is a 1000-hour multilingual South African speech dataset spanning 8 languages, designed for automatic speech recognition (ASR), code-switching research, and large audio model (LAM) evaluation.
This repository provides a 69-minute evaluation preview (Episode 001 — Nobantu Vilakazi). The preview mirrors the segmentation, metadata schema, validation pipeline, and structural standards used in the full… See the full description on the dataset page: https://huggingface.co/datasets/ODYSSEYAILABS/odyssey-sa-voice-v0.4-sample.
