datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
5bfe2d098c8486d97fac8be76d86ec9146435245
train
56:46:32
50,557
589,095
11.7
31.9
techiaith/corpws-clllc-wlga
5d00294c31c78b1d7937bb2c2bc6cc70bc18d410
clips
48:20:49
27,579… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2607.spite-CV16-TP9B
Spite Dataset
Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from Tower-Plus-9B.
Configs
en_de
en_es
en_fr
en_it
en_ko
en_nl
en_pt
en_ru
en_zh
Usage
from datasets import load_dataset
ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt")
spite-CV16-Euro9B
Spite Dataset
Pseudolabeled speech translation data with quality annotations from multiple metrics. This version uses transcripts from Common Voice 16.1 and translations from EuroLLM-9B-Instruct.
Configs
en_de
en_es
en_fr
en_it
en_ko
en_nl
en_pt
en_ru
en_zh
Usage
from datasets import load_dataset
ds = load_dataset("bpop/spite-CV16-Euro9B", "en_pt")
preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
28.2
techiaith/corpws-clllc-wlga
2603_rc3
clips
41:54:21
29,446
447,780
15.2
22.4
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2603.preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2605
main
train
57:14:31
50,934
594,564
11.7
29.5
techiaith/corpws-clllc-wlga
main
clips
52:27:50
29,794
544,066
18.3
27.1
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2606.preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2601
main
train
52:38:17
48,608
573,772
11.8
27.7
techiaith/corpws-clllc-wlga
main
clips
20:00:52
18,905
228,962
12.1
10.5
techiaith/commonvoice_23_0_cy
main… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-cven-wlga-ca-ec-2601.preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Dataset Card
Preprocessed Dataset: DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602
Revision: main
Dataset Statistics
Train Split Statistics
Dataset
Revision
Split
Duration (HH:MM:SS)
Clips
Words
Words/Clip
%
DewiBrynJones/banc-trawsgrifiadau-bangor-2602
main
train
52:45:22
48,569
556,542
11.5
32.2
techiaith/corpws-clllc-wlga
main
clips
20:02:39
18,851
228,230
12.1
12.2
techiaith/commonvoice_23_0_cy
main
train+dev+other_with_excluded… See the full description on the dataset page: https://huggingface.co/datasets/DewiBrynJones/preprocessed-whisper-btb-cv-cvad-wlga-ca-2602.
