Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mort666 /cv_corpus_v22 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. NOTE: currently converting to parquet for convenience.. WIP Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.audioautomatic-speech-recognition1M<n<10M0 likes5k downloads10mo agoHugging Face02changelinglab /cv-v1.0-segment CommonVoice v1 Phone-Segment Alignments Phone-level time alignments for 10 languages of Mozilla Common Voice, packaged in a canonical segmentation schema with embedded 16 kHz audio. The phone boundaries come from the charsiu/cv_ali release of MFA alignments; the audio and transcripts come from Common Voice Corpus 13.0 (2023-03-09). Dataset summary lang train rows train hrs val rows val hrs test rows test hrs en 1,008,669 1,354.0 3,537 4.9 1,285 1.7 rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.audioautomatic-speech-recognition1M<n<10M3 likes1.9k downloads6mo agoHugging Face03FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes674 downloads6mo agoHugging Face04fidoriel /cv-22-deGerman split of Common Voice 22. cc0 license audioautomatic-speech-recognition100K<n<1M3 likes511 downloads1y agoHugging Face05masuidrive /cv-corpus-17.0-zh-CN-client_id-grouped cv-corpus-17.0-zh-CN-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-CN-client_id-grouped.audioautomatic-speech-recognition100K<n<1M3 likes475 downloads2y agoHugging Face06LokaalHub /nl-asr-cv Dutch ASR (Common Voice, speaker-disjoint splits) Dutch (nl) speech for ASR, built from Mozilla Common Voice (CC0) via the open fsicoli/common_voice_17_0 mirror. Built to fine-tune tiny ASR models (e.g. openai/whisper-tiny). Splits Split Hours train 94.4 dev 0.8 test 2.0 Held-out dev/test are disjoint from train by both speaker and sentence. Common Voice's official dev/test are capped by whole speakers (dev ~0.75h, test ~2.0h) with the… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nl-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes451 downloads4mo agoHugging Face07masuidrive /cv-corpus-17.0-zh-TW-client_id-grouped cv-corpus-17.0-zh-TW-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-zh-TW-client_id-grouped.audioautomatic-speech-recognition10K<n<100K1 likes384 downloads2y agoHugging Face08LokaalHub /nb-NO-asr-cv Norwegian Bokmål ASR (Common Voice 22, filtered + rebalanced) Norwegian Bokmål (nb-NO) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open NbAiLab/NPSC mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 44302 86.9 dev 453 0.9 test 1211 2.2 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/nb-NO-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes299 downloads4mo agoHugging Face09LokaalHub /cy-asr-cv Welsh ASR (Common Voice 22, filtered + rebalanced) Welsh (cy) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train ? 50.1 dev ? 0.8 test ? 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's official… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/cy-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes263 downloads4mo agoHugging Face10speech-uk /cv22-opus Common Voice for 🇺🇦 Ukrainian (OPUS) Ukrainian validated subset of Common Voice 22 Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 89248 Total duration: 115h 5m 9s audioautomatic-speech-recognition10K<n<100K0 likes239 downloads6mo agoHugging Face11LokaalHub /frisian-asr-cv22 Frisian ASR (Common Voice 22, filtered) Open Standard West Frisian (fy-NL) speech for ASR, built from Mozilla Common Voice 22.0 (CC0). The validated training split is augmented with the unvalidated other bucket, which is auto-filtered by CTC agreement with the known prompt using a Frisian-specialized wav2vec2 model. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours Composition train 29,929… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/frisian-asr-cv22.audioautomatic-speech-recognition10K<n<100K0 likes205 downloads4mo agoHugging Face12masuidrive /cv-corpus-17.0-ja-client_id-grouped cv-corpus-17.0-ja-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 30 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-17.0-ja-client_id-grouped.audioautomatic-speech-recognition10K<n<100K2 likes158 downloads2y agoHugging Face13LokaalHub /sv-SE-asr-cv Swedish ASR (Common Voice 22, filtered + rebalanced) Swedish (sv-SE) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 22166 26.1 dev 694 0.8 test 1602 2.0 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/sv-SE-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes123 downloads4mo agoHugging Face14DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 5bfe2d098c8486d97fac8be76d86ec9146435245 train 56:46:32 50,557 589,095 11.7 31.9 techiaith/corpws-clllc-wlga 5d00294c31c78b1d7937bb2c2bc6cc70bc18d410 clips 48:20:49 27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.tabularautomatic-speech-recognition100K<n<1M0 likes112 downloads2mo agoHugging Face15Elyadata /CV18-NER CV-18 NER CV-18 NER is the first publicly available dataset for Named Entity Recognition (NER) from Arabic speech. It was created by augmenting the Arabic Common Voice 18 corpus with manual NER annotations following the fine-grained Wojood schema, which covers 21 entity types. The dataset provides a benchmark for evaluating both pipeline systems (ASR + text NER) and end-to-end speech NER models. It is particularly valuable for research in low-resource settings and morphologically… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/CV18-NER.automatic-speech-recognition1 likes101 downloads5mo agoHugging Face16masuidrive /cv-corpus-1.0-en-client_id-grouped cv-corpus-1.0-en-client_id-grouped This dataset is a subset of the Common Voice dataset, filtered and grouped based on the client ID (treated as speaker ID). Dataset Details The dataset is derived from the Common Voice dataset. The original dataset is available at Common Voice Dataset. The dataset is grouped by client ID, which is treated as the speaker ID for this dataset. Each group is filtered to include only client IDs with a minimum of 60 samples and a maximum of… See the full description on the dataset page: https://huggingface.co/datasets/masuidrive/cv-corpus-1.0-en-client_id-grouped.audioautomatic-speech-recognition100K<n<1M1 likes85 downloads2y agoHugging Face17LokaalHub /da-asr-cv Danish ASR (Common Voice 22, filtered + rebalanced) Danish (da) speech for ASR, built from Mozilla Common Voice 22.0 (CC0) via the open fsicoli/common_voice_22_0 mirror. Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b). Splits Split Clips Hours train 8270 10.0 dev 734 0.9 test 1593 2.1 train = the official validated train split + the filtered other bucket + the excess dev/test speakers: Common Voice's… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/da-asr-cv.audioautomatic-speech-recognition10K<n<100K0 likes84 downloads4mo agoHugging Face18projecte-aina /cv17_es_other_automatically_verifiedSplit called -other- of the Spanish Common Voice v17.0 that was automatically verified using various ASR system.automatic-speech-recognition100K<n<1M2 likes73 downloads1y agoHugging Face19Porameht /processed-cv-17-th-130k processed-cv-17-th-130k Cleaned Thai split of Mozilla Common Voice 17: 130,551 utterances (117,536 train / 3,950 dev / 9,065 test) with transcripts, ready for ASR training. Format Field Description sentence Transcript in Thai audio Audio clip Usage from datasets import load_dataset ds = load_dataset("Porameht/processed-cv-17-th-130k") Source and license Derived from Mozilla Common Voice 17.0 (Thai), released under… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/processed-cv-17-th-130k.audioautomatic-speech-recognition100K<n<1M1 likes66 downloads1mo agoHugging Face20Yehor /cv10-uk-testset-clean The cleaned Common Voice 10 (test set) that has been checked by a human for Ukrainian 🇺🇦 Overview This repository contains the archive of Common Voice 10 (test set) with checked Ukrainian transcriptions and audios. All audios have been checked by a human to be sure that they are correct. This archive is used to test all ASR models listed here: https://github.com/egorsmkv/speech-recognition-uk Community Discord: https://bit.ly/discord-uds Speech… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/cv10-uk-testset-clean.audioautomatic-speech-recognition1K<n<10K3 likes61 downloads2y agoHugging Face21bpop /spite-CV16-TP9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from Tower-Plus-9B. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes39 downloads8mo agoHugging Face22bilguun /cv-mn-24.0 cv-mn-24.0 Mongolian subset of the Mozilla Common Voice speech recognition dataset. Dataset Statistics Total samples: 6,018Total duration: 9h 7m 2s (9.12 h) Per-split breakdown Split Samples Total Duration Avg Duration train 2,188 3h 7m 34s (3.13 h) 5.14 s validation 1,896 2h 54m 41s (2.91 h) 5.53 s test 1,934 3h 4m 47s (3.08 h) 5.73 s audioautomatic-speech-recognition1K<n<10K1 likes37 downloads6mo agoHugging Face23Trelis /cv-en-scripted-test-500 Common Voice English Scripted Test Set — 500 clips n = 500 utterances · private eval set for ASR benchmarking Source Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball). Construction Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.audioautomatic-speech-recognitionn<1K0 likes35 downloads5mo agoHugging Face24Veronica1NW /cv17_sw_kenyan_sample Common Voice 17.0 — Swahili (Kenyan Sample) This dataset is a filtered sample of the Mozilla Common Voice 17.0 corpus, focusing on Swahili (sw) speech with Kenyan voices. It has been subsetted for experimentation and prototyping in ASR (Automatic Speech Recognition) models targeting speech-impaired users in Kenya, covering Kenyan English and Kiswahili. Dataset Summary Language: Kiswahili (Swahili, sw) Accent/Region: Kenyan speakers Domain: Conversational… See the full description on the dataset page: https://huggingface.co/datasets/Veronica1NW/cv17_sw_kenyan_sample.audioautomatic-speech-recognition1K<n<10K0 likes34 downloads1y agoHugging Face25tiny-aya-translate /cv-tr-eval Common Voice Turkish Eval 4,825 Turkish test clips (~45 MB, 16 kHz) in the Mozilla Common Voice schema: transcription, duration, up_votes / down_votes, and the age / gender / accent speaker attributes. Schema in the YAML header above. An evaluation-only Turkish counterpart to lahaja-eval; never trained on. Used to sanity-check Turkish ASR quality on real human speech, which matters here because the v0.3 training corpus is entirely synthetic TTS and the released model is… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/cv-tr-eval.audioautomatic-speech-recognition1K<n<10K0 likes32 downloads3mo agoHugging Face26bpop /spite-CV16-Euro9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from EuroLLM-9B-Instruct. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes30 downloads8mo agoHugging Face27TriLingDATA /si-ta-en-asr-cv-debug Sinhala-Tamil-English ASR Dataset A combined speech recognition dataset for Sinhala, Tamil and English, built for fine-tuning multilingual ASR models such as Whisper. All audio is 16 kHz mono 16-bit WAV, transcripts keep cleaned punctuation and casing, and test/validation sets are held out by speaker. Tamil comes from two sources that can be selected with the source column: indic_tts_tamil (IndicTTS, spontaneous talk-style speech) and common_voice_22.0_tamil (Common Voice 22.0… See the full description on the dataset page: https://huggingface.co/datasets/TriLingDATA/si-ta-en-asr-cv-debug.audioautomatic-speech-recognition1K<n<10K0 likes29 downloads2d agoHugging Face28cvxhull /stt-calibration STT Calibration Dataset Tiny calibration dataset for PersonalAssistant STT service. Used on first run to auto-tune speculative pre-transcription parameters. Contents File Duration Size Purpose short.wav 3.5s 110KB RTF measurement + VAD onset latency long.wav 23.3s 729KB Split quality calibration (whole vs split comparison) very_long.wav 56.8s 1.8MB Multi-split calibration (find minimum safe split interval) manifest.json - 2KB Sample metadata + reference… See the full description on the dataset page: https://huggingface.co/datasets/cvxhull/stt-calibration.audioautomatic-speech-recognitionn<1K0 likes22 downloads7mo agoHugging Face29instinct-org /cv_chunked_speech_restorisedgated cv_chunked_speech_restorised This is a gated Uzbek speech-restorised chunked speech dataset from instinct-org. This repository contains speech audio and transcripts for speech-to-text training, evaluation, or data preparation workflows. Language Primary language: uz (Uzbek) Intended Use speech-to-text training and evaluation Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_speech_restorised.audioautomatic-speech-recognition10K<n<100K0 likes17 downloads2mo agoHugging Face30speech-uk /cv22gated Common Voice for 🇺🇦 Ukrainian Ukrainian validated subset of Common Voice 22 Community Discord: https://bit.ly/discord-uds Speech Recognition: https://t.me/speech_recognition_uk Speech Synthesis: https://t.me/speech_synthesis_uk Stats Total files processed: 89248 Total duration: 115h 5m 9s audioautomatic-speech-recognition10K<n<100K0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.