Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes674 downloads6mo agoHugging Face02DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 5bfe2d098c8486d97fac8be76d86ec9146435245 train 56:46:32 50,557 589,095 11.7 31.9 techiaith/corpws-clllc-wlga 5d00294c31c78b1d7937bb2c2bc6cc70bc18d410 clips 48:20:49 27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.tabularautomatic-speech-recognition100K<n<1M0 likes112 downloads2mo agoHugging Face03bpop /spite-CV16-TP9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from Tower-Plus-9B. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes39 downloads8mo agoHugging Face04bpop /spite-CV16-Euro9B Spite Dataset Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from EuroLLM-9B-Instruct. Configs en_de en_es en_fr en_it en_ko en_nl en_pt en_ru en_zh Usage from datasets import load_dataset ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt") tabulartranslation1M<n<10M0 likes30 downloads8mo agoHugging Face05DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2603gated Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2602 main train 52:45:22 48,569 556,542 11.5 28.2 techiaith/corpws-clllc-wlga 2603_rc3 clips 41:54:21 29,446 447,780 15.2 22.4 techiaith/commonvoice_23_0_cy main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603.tabularautomatic-speech-recognition100K<n<1M0 likes11 downloads7mo agoHugging Face06DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2606gated Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2605 main train 57:14:31 50,934 594,564 11.7 29.5 techiaith/corpws-clllc-wlga main clips 52:27:50 29,794 544,066 18.3 27.1 techiaith/commonvoice_23_0_cy main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606.tabularautomatic-speech-recognition100K<n<1M0 likes10 downloads4mo agoHugging Face07DewiBrynJones /preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601gated Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2601 main train 52:38:17 48,608 573,772 11.8 27.7 techiaith/corpws-clllc-wlga main clips 20:00:52 18,905 228,962 12.1 10.5 techiaith/commonvoice_23_0_cy main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601.tabularautomatic-speech-recognition100K<n<1M0 likes9 downloads8mo agoHugging Face08DewiBrynJones /preprocessed-whisper-btb-cv-cvad-wlga-ca-2602gated Dataset Card Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602 Revision: main Dataset Statistics Train Split Statistics Dataset Revision Split Duration (HH:MM:SS) Clips Words Words/Clip % DewiBrynJones/banc-trawsgrifiadau-bangor-2602 main train 52:45:22 48,569 556,542 11.5 32.2 techiaith/corpws-clllc-wlga main clips 20:02:39 18,851 228,230 12.1 12.2 techiaith/commonvoice_23_0_cy main train+dev+other_with_excluded… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602.tabularautomatic-speech-recognition100K<n<1M0 likes7 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.