datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kare-codeswitch-samples
Kare — Code-Switching Illustrative Samples
Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin
code-switching that Kare, a voice-first
AI health assistant for Nigeria, is built to understand — submitted as part
of Kare's entry to the Sahara CodeSwitch Africa Challenge.
What this is — and isn't
Is: eight original sentences, written by the Kare team specifically
for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian
Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.JA_audio_JA_text_180k_samples-Noise and silence have been removed from the begining and end of each sample.
-Unnecessary and inaccurate punctuation have been removed.
-Text has been normalized.
Normalization is based on the neologd's rules: https://github.com/neologd/mecab-ipadic-neologd/wiki/Regexp.ja.
clinical-speech-samples
🏥 Clinical Speech Samples Dataset
High-Quality Clinical Audio Dataset for Medical AI Research
Curated dataset for training and evaluating clinical speech processing models
📖 Dataset Description
This dataset contains de-identified clinical audio samples for research and development of medical speech processing systems. It's designed to support training and evaluation of:
🎤 Speech enhancement in clinical settings
🔊 Audio source separation for medical… See the full description on the dataset page: https://huggingface.co/datasets/bhaskarvilles/clinical-speech-samples.luel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.muscat-merged-samples
MUSCAT — Merged Long-Form Samples
This dataset is a merged, long-form reformatting of
goodpiku/muscat-eval
(MUSCAT: A Multi-Device Dataset for Code-Switching ASR and Segmentation
Evaluation).
The original MUSCAT release stores each conversation as many short,
single-language segments. Here those segments are concatenated back into one
continuous recording per conversation, so each row is a single long-form
code-switching audio with inline language/timing markers. The layout… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/muscat-merged-samples.mandarin-speech-samples
Mandarin Speech Samples
This sample shows Mandarin Chinese speech with clip-level metadata and preview transcripts. It is meant to help buyers review language fit, recording quality, and sample structure before scoping a larger delivery.
What This Shows
Mandarin speech audio with consistent metadata
Clip-level transcript fields for content review
Language and format signals for procurement review
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/mandarin-speech-samples.english-x-code-switching-samples
Synthetic English Code-Switching Evaluation Set Samples
This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset.
Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row.
Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.bengali-multi-speaker-speech-samples
Bengali Speech: Multi-Speaker Samples
This sample shows Bengali multi-speaker speech with aligned ground-truth transcripts. It is meant to help buyers review conversational structure, speaker overlap, transcript quality, and audio consistency before scoping a larger delivery.
What This Shows
Multi-speaker Bengali speech with transcript alignment
Conversation-style audio rather than isolated prompt reading
Metadata that distinguishes language, format, and speaker… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-multi-speaker-speech-samples.korean-speech-samples
Korean Speech Samples
This sample shows Korean contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Korean speech recordings from contributor collection workflows
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/korean-speech-samples.french-speech-samples
French Speech Samples
This sample shows French contributor speech paired with source transcripts. It is meant to help buyers review recording quality, transcript alignment, and metadata structure before scoping a larger delivery.
What This Shows
French single-speaker recordings
Transcript alignment from the source dataset
Clip-level metadata for format and review context
Dataset Specifications
Field
Value
Modality
Audio
Language… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/french-speech-samples.telugu-speech-samples
Telugu Speech Samples
This sample shows Telugu speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery.
What This Shows
Telugu speech recordings with paired transcripts
Ground-truth labels at the clip level
Format metadata for review and delivery planning
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/telugu-speech-samples.cantonese-speech-samples
Cantonese Speech Samples
This sample shows Cantonese speech with native transcript metadata. It is meant to help buyers review dialect fit, recording quality, and sample structure before requesting broader coverage.
What This Shows
Cantonese speech audio with paired transcript metadata
Language-specific metadata for review and delivery planning
A compact preview of the available sample structure
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/cantonese-speech-samples.tamil-speech-samples
Tamil Speech Samples
This sample shows Tamil speech with ground-truth transcripts and consistent audio metadata. It is meant to help buyers review language fit, transcript quality, and capture format before scoping a larger delivery.
What This Shows
Tamil speech recordings with paired transcripts
Ground-truth labels at the clip level
Format metadata for review and delivery planning
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/tamil-speech-samples.english-en-x-code-switching-main-lang-samples
English EN-X Code-Switching Main-Language Samples
This dataset contains the individual full FLEURS utterance chunks used to build the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples.english-en-x-code-switching-main-lang-samples-merged
English EN-X Code-Switching Main-Language Merged Samples
This dataset contains contiguous same-language segments from the paired mixed dataset.
Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly.
Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang-samples-merged.hindi-speech-samples
Hindi Speech Samples
This sample shows Hindi contributor speech paired with validated text. It is meant to help buyers review spoken content, transcript alignment, and audio format consistency before scoping a larger delivery.
What This Shows
Single-speaker Hindi recordings from contributor collection workflows
Ground-truth transcript alignment at the clip level
Audio metadata suitable for evaluating format and capture consistency
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/hindi-speech-samples.henan-speech-samples
Henan Speech Samples
This sample shows Henan Chinese speech with native transcript metadata. It is meant to help buyers review regional speech fit, recording quality, and sample structure before requesting broader coverage.
What This Shows
Henan Chinese speech audio with paired transcript metadata
Regional language metadata for review and delivery planning
A compact preview of the available sample structure
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/henan-speech-samples.urdu-speech-samples
Urdu Speech Samples
This sample shows Urdu speech with transcript alignment and simple audio metadata. It is meant to help buyers review language fit and capture quality before requesting a larger sample or production delivery.
What This Shows
Urdu speech audio with paired text
A compact view of transcript and metadata structure
Audio format signals for procurement review
Dataset Specifications
Field
Value
Modality
Audio
Language
Urdu… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/urdu-speech-samples.spanish-speech-samples
Spanish Speech Samples
This sample shows Spanish contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Spanish speech samples with clip-level review metadata
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field
Value… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/spanish-speech-samples.bengali-speech-samples
Bengali Speech Samples
This sample shows Bengali read and conversational speech with paired transcripts. It is meant to help buyers review spoken content, transcript alignment, and audio consistency before scoping a larger delivery.
What This Shows
Bengali speech across read and conversational styles
Clip-level transcript alignment
Audio metadata that supports format and quality review
Dataset Specifications
Field
Value
Modality
Audio… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/bengali-speech-samples.Unique007_french_phone_1421_samples_16khzCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données Unique007/french_phone_1421_samples_16khz.
speech-transcription-samples
Speech Transcription Samples Dataset
Synthetic speech-to-text transcription samples for ASR research.
