Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bandad /asr-testset-kw-ja-v1 日本語 ASR アノテーション v1 重要:評価結果を報告する際の規約 本テストセットで評価結果を報告する際は、事前学習を含む学習にYouTubeの音声を使用したかどうかと、次の評価区分を必ず明記してください。 使用した場合:in-domain評価 使用していない場合:out-of-domain評価 この区分は、音声認識の評価結果を比較する上で重要です。最終的な追加学習だけでなく、使用するモデルの事前学習も含めて判断してください。 元データセットの音声に、人手で区間ごとの転記・タグ・採否を付けたデータです。音声の内容、ディレクトリ構成、ファイル名は元のままです。 ファイルと表示 **annotations.jsonl**:提出済みの全結果。1行が1音声です。 **metadata.jsonl**:HF表示用に自動生成したデータ。音声全体が不使用の行を除き、audio を file_name に置き換えています。 **data/**:採用した音声ファイル。 HFのビューアーは… See the full description on the dataset page: https://huggingface.co/datasets/bandad/asr-testset-kw-ja-v1.audioautomatic-speech-recognitionn<1K0 likes494 downloads6d agoHugging Face02ggfox00000 /stt-vibravox-fr-test VibraVox FR — test split (mirror of Cnam-LMSSC/vibravox) Mirror public des splits test de VibraVox (CNAM-LMSSC, Paris) pour benchmark ASR français multi-capteur sur audio standard ET non-standard (bone-conduction, in-ear, throat, accéléromètre). Ce repo contient uniquement les configs speech_clean + speech_noisy (les seules avec transcription). Les configs speechless_* upstream sont exclues car sans texte → pas de WER possible. Configs Config Test rows Test… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-vibravox-fr-test.audioautomatic-speech-recognition1K<n<10K0 likes377 downloads6mo agoHugging Face03yuekai /speechio_test SpeechIO ASR Test Sets (parquet) Parquet repackaging of the SpeechColab SpeechIO Mandarin ASR benchmark, re-exported from yuekai/speechio (Lhotse cuts) into standard HuggingFace parquet with embedded 16 kHz audio. 27 test sets: SPEECHIO_ASR_ZH00000 ... SPEECHIO_ASR_ZH00026, each a config with a single test split. ~43k utterances, ~66 hours total, evaluation only. Columns column type note segment_id string utterance id speaker string speaker id… See the full description on the dataset page: https://huggingface.co/datasets/yuekai/speechio_test.audioautomatic-speech-recognition10K<n<100K0 likes287 downloads3mo agoHugging Face04rasgaard /fleurs_test FLEURS Test Dataset with Enhanced Metadata This dataset is an enhanced version of the FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech) test set, restructured with complete metadata for easier use in automatic speech recognition (ASR) and multilingual speech processing tasks. Dataset Description FLEURS is a multilingual speech benchmark dataset designed to evaluate universal speech representations. This particular version focuses on 25 European… See the full description on the dataset page: https://huggingface.co/datasets/rasgaard/fleurs_test.audioautomatic-speech-recognition10K<n<100K0 likes227 downloads8mo agoHugging Face05humair025 /h-test@misc{humair025/h-test, title = {h-test}, author = {Humair Munir}, year = {2025}, howpublished = {\url{https://huggingface.co/datasets/humair025/h-test}}, note = {Synthetic dataset of speech (with emotions). Licensed under CC-BY 4.0.} } audiotext-to-speechn<1K0 likes200 downloads10mo agoHugging Face06ciempiess /ciempiess_testThe CIEMPIESS TEST Corpus is a gender balanced corpus destined to test acoustic models for the speech recognition task. The corpus was manually transcribed and it contains audio recordings from 10 male and 10 female speakers. The CIEMPIESS TEST is one of the three corpora included at the LDC's \"CIEMPIESS Experimentation\" (LDC2019S07).audioautomatic-speech-recognition1K<n<10K3 likes175 downloads3y agoHugging Face07shrikanth-19 /dhravani-mit-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-mit-test.audioautomatic-speech-recognitionn<1K0 likes172 downloads2d agoHugging Face08FBK-MT /Speech-MASSIVE-testgated Speech-MASSIVE Test Split This dataset repository is only for test split of Speech-MASSIVE. train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE Dataset Description Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.audioaudio-classification10K<n<100K9 likes165 downloads1y agoHugging Face09shrikanth-19 /dhravani-iitpatna-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitpatna-test.audioautomatic-speech-recognitionn<1K0 likes147 downloads2d agoHugging Face10shrikanth-19 /dhravani-IIT_Guwahati-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IIT_Guwahati-test.audioautomatic-speech-recognitionn<1K0 likes146 downloads2d agoHugging Face11shrikanth-19 /dhravani-iitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes134 downloads2d agoHugging Face12shrikanth-19 /dhravani-IGDTUW_Delhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-IGDTUW_Delhi-test.audioautomatic-speech-recognitionn<1K0 likes131 downloads2d agoHugging Face13shrikanth-19 /dhravani-iiitdelhi-testCheck out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference Dataset Preparation Interface for Fine-tuning Whisper A web-based interface for preparing audio datasets to fine-tune OpenAI's Whisper model. This tool helps in recording, managing, and organizing voice recordings with their corresponding transcriptions, with support for cloud storage and authentication. Features 🔐 User authentication via Pocketbase ☁️ Cloud storage… See the full description on the dataset page: https://huggingface.co/datasets/shrikanth-19/dhravani-iiitdelhi-test.audioautomatic-speech-recognitionn<1K0 likes130 downloads2d agoHugging Face14ggfox00000 /stt-cefc-fr-test CEFC-Orfeo FR — long-form oral test mirror Mirror non-officiel du Corpus d'Études du Français Contemporain (CEFC) agrégé par le projet Orfeo (ANR), tel que distribué sur le portail projet-orfeo.fr (release 13). Long-form : 1 row = 1 fichier audio entier (30-60 min en moyenne). 12 sous-corpus oraux du français contemporain, 303 heures au total, 901 fichiers. Idéal pour bench Whisper / Canary en conditions réelles (chunked decoding, dérive temporelle, multi-locuteurs).… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-cefc-fr-test.audioautomatic-speech-recognitionn<1K0 likes125 downloads6mo agoHugging Face15Yehor /RS-test-fix2audioautomatic-speech-recognition10K<n<100K0 likes113 downloads3mo agoHugging Face16NVVSpeech-Challenge /NVVSpeech-Challenge-Track1-Test-Set NVVSpeech Challenge Track 1 Test Set Track 1 test set for the NVVSpeech Challenge at ISCSLP 2026. Task Given a speech recording, produce a transcript that contains the spoken content and the non-verbal vocalization (NVV) tags at their corresponding positions. Dataset Summary Language Samples Chinese 985 English 961 Total 1,946 Files . ├── README.md ├── SUBMISSION_GUIDE.txt ├── test.jsonl ├── ground_truth.jsonl ├──… See the full description on the dataset page: https://huggingface.co/datasets/NVVSpeech-Challenge/NVVSpeech-Challenge-Track1-Test-Set.audioautomatic-speech-recognition1K<n<10K0 likes106 downloads12d agoHugging Face17anjalyjayakrishnan /testThe Snow Mountain dataset contains the audio recordings (in .mp3 format) and the corresponding text of The Bible in 11 Indian languages. The recordings were done in a studio setting by native speakers. Each language has a single speaker in the dataset. Most of these languages are geographically concentrated in the Northern part of India around the state of Himachal Pradesh. Being related to Hindi they all use the Devanagari script for transcription.audioautomatic-speech-recognition1K<n<10K0 likes100 downloads4y agoHugging Face18Yehor /RS-testaudioautomatic-speech-recognition10K<n<100K0 likes96 downloads3mo agoHugging Face19ggfox00000 /stt-summre-fr-test SUMM-RE — French test split (mirror of linagora/SUMM-RE) Mirror public du split test de SUMM-RE (LINAGORA / Aix-Marseille LPL), pour benchmark ASR français conversationnel (parole de réunion, 3-4 locuteurs, ~20 min par session). ⚠ Ce repo ne contient que le split test (124 tracks individuelles = 37 réunions × 3-4 micros). Pour les splits train / dev, voir le repo upstream linagora/SUMM-RE. Contenu 124 pistes audio individuelles (1 piste = 1 microphone d'un locuteur… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-summre-fr-test.audioautomatic-speech-recognitionn<1K0 likes93 downloads6mo agoHugging Face20XRXRX /X-Voice-TestsetX-Voice Multilingual Test Set High-Fidelity Test Set for Multilingual Text-to-Speech across 30 Languages This test set is built as part of the research: X-Voice: One Speaker, 30+ Languages with Zero-Shot Voice Cloning, serving as the evaluation benchmark for our model. Dataset Summary 30 languages European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi (Finnish), fr (French), hr (Croatian), hu (Hungarian), it… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Testset.audiotext-to-speech10K<n<100K4 likes87 downloads5mo agoHugging Face21freyavoice /common-voice-17-tr-test common-voice-17-tr-test Turkish test split of Common Voice 17.0 (tr), re-hosted for Turkish STT benchmarking. Rows: 11290 Columns: client_id, path, audio, sentence, up_votes, down_votes, age, gender, accent, locale, segment, variant Source: https://commonvoice.mozilla.org License: cc0-1.0 (inherited from source) Only the Turkish test split is included, extracted as-is from the source dataset. audioautomatic-speech-recognition10K<n<100K0 likes86 downloads4mo agoHugging Face22JacobLinCool /audio-testing audio-testing Overview This is a small, open dataset designed for quick validation of audio-related pipelines and applications, especially for Text-to-Speech (TTS) and Speech-to-Text (STT) systems. It provides a few short, diverse audio clips and corresponding text transcripts, allowing developers to verify input/output handling, audio processing, and transcription logic without downloading large datasets. Contents 3 short audio samples (.mp3, .wav)… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/audio-testing.audioautomatic-speech-recognitionn<1K0 likes85 downloads1y agoHugging Face23adalat-ai /vividh-test-hindi 🎙️ Vividh-ASR Benchmark — Hindi (Test Split) How well does your ASR model actually work in the wild? Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on read… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-hindi.audioautomatic-speech-recognition10K<n<100K1 likes82 downloads5mo agoHugging Face24xiaofff /omnievalkit-data-test OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 89 Total samples: 353,610 Total size: 352.3 GB (Parquet with embedded audio/image, no video) Subsets requiring video download: 42 Note: Video files are NOT embedded in the Parquet files due to size constraints. Usage from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/xiaofff/omnievalkit-data-test.audioaudio-classification10K<n<100K0 likes78 downloads7mo agoHugging Face25Trelis /ami-2speaker-test AMI 2-Speaker Test Set Need a voice model for your domain? Trelis builds custom ASR, TTS, and voice agent pipelines for specialist verticals (legal, medical, finance, construction) and low-resource languages. Enquire or book a consultation → A 50-clip benchmark for 2-speaker overlapping speech recognition, derived from the AMI Meeting Corpus test split. Each clip is 8–28 seconds of real conversational meeting audio reconstructed as a 2-speaker virtual meeting, with separate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/ami-2speaker-test.audioautomatic-speech-recognitionn<1K0 likes78 downloads6mo agoHugging Face26adalat-ai /vividh-test-malayalam 🎙️ Vividh-ASR Benchmark — Malayalam (Test Split) How well does your ASR model actually work in the wild?Vividh-ASR is a complexity-stratified benchmark that tells you exactly where your model succeeds — and where it falls apart. Most Indic ASR benchmarks evaluate models on clean, studio-recorded speech. Real-world audio is not that. Vividh-ASR organises evaluation by acoustic complexity rather than domain, exposing the studio-bias that plagues models fine-tuned predominantly on… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/vividh-test-malayalam.audioautomatic-speech-recognition10K<n<100K1 likes76 downloads5mo agoHugging Face27orkidea /wayuu_CO_test Dataset Audio Duration The dataset consists of 810 audio recordings, each accompanied by its respective transcription. The lexical corpus encompasses approximately 1,000 unique words. Total Audio Duration: 2801 seconds (approximately 34 minutes) Average Audio Duration: 3.41 seconds The dataset offers valuable insights into the Wayuunaiki language's phonetic and linguistic characteristics. It's important to note that the dataset originates from recordings and transcriptions of the… See the full description on the dataset page: https://huggingface.co/datasets/orkidea/wayuu_CO_test.audioautomatic-speech-recognitionn<1K2 likes71 downloads3y agoHugging Face28sfsmcnulty /stt_test_audio Mobile Voice Platform STT Test Audio Versioned benchmark audio for the MobileVoicePlatform-Android SampleApp. The repository intentionally contains two views of the same clips: data/ is an AudioFolder-compatible view with metadata.csv for Hugging Face tooling. packs/ contains checksum-pinned ZIPs optimized for bounded download and validation on Android. catalog.json is the machine-readable index used to discover datasets, categories, languages, checksums, clip references, and… See the full description on the dataset page: https://huggingface.co/datasets/sfsmcnulty/stt_test_audio.audioautomatic-speech-recognitionn<1K0 likes70 downloads1mo agoHugging Face29arda-argmax /fastmss-v0.5.0-test FastMSS synthetic multi-speaker meetings - parquet edition Streaming-friendly parquet shards of the FastMSS synthetic multi-speaker conversational corpus. Each row is one mixture with the audio bytes embedded inline (16 kHz mono WAV) plus per-segment diarization timestamps, per-word transcript and the full lhotse cut as a JSON blob. See fastmss/hf_dataset.py for the schema docstring. Subsets and splits v0.5.0_test — splits: train — 1000 mixtures, 1081.0 min total, 3609… See the full description on the dataset page: https://huggingface.co/datasets/arda-argmax/fastmss-v0.5.0-test.audioautomatic-speech-recognition1K<n<10K0 likes67 downloads5mo agoHugging Face30test12313 /Japanese-Eroge-Voice Japanese-Eroge-Voice Description This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model. Preprocessing Steps The raw audio data has undergone the following preprocessing steps: Loudness Normalization: Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/test12313/Japanese-Eroge-Voice.audiotext-to-speech100K<n<1M1 likes63 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.