datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
riksdagen_testsamromur_children_test
Dataset Card for samromur_children
Dataset Summary
The Samrómur Children Corpus consists of audio recordings and metadata files containing prompts read by the participants. It contains more than 137000 validated speech-recordings uttered by Icelandic children.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab (LVL) at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology. The recording process has… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/samromur_children_test.Lok_Sabha_test
Lok Sabha Spoken English Corpus — ParlaSpeech-compatible pilot
This eight-hour pilot follows the Hugging Face structure used by ParlaSpeech-style
speech corpora, with a compact schema tailored to Lok Sabha data.
The default configuration contains one accepted aligned audio segment per row,
with embedded 16 kHz audio, verbatim ASR, an explicitly separate edited UCR
passage, word timings, speaker metadata, and source-order fields.
Configurations
default: 1,195… See the full description on the dataset page: https://huggingface.co/datasets/stcoats/Lok_Sabha_test.youtube_ua_noisy_subtitles_test
The list of all subsets in the dataset
Each subset is generated splitting videos from given particular ukrainiam YouTube channel
All subsets are in test split
"opodcast" subset is from channel "О! ПОДКАСТ"
"rozdympodcast" subset is from channel "Роздум | Подкаст"
"test" subset is just a small subset of samples
Loading a particular subset
>>> data_files = {"train": "data/<your_subset>.parquet"}
>>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_noisy_subtitles_test.youtube_ua_subtitles_test
The list of all subsets in the dataset
Each subset is generated splitting videos from given particular ukrainiam YouTube channel
All subsets are in test split
"opodcast" subset is from channel "О! ПОДКАСТ"
"rozdympodcast" subset is from channel "Роздум | Подкаст"
"test" subset is just a small subset of samples
Loading a particular subset
>>> data_files = {"train": "data/<your_subset>.parquet"}
>>> data = load_dataset("Zarakun/youtube_ua_subtitles_test"… See the full description on the dataset page: https://huggingface.co/datasets/Zarakun/youtube_ua_subtitles_test.
