datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shkolkovo-bobr.video-webinars-audio
shkolkovo-bobr.video-webinars-audio
Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo.
Language: Russian, includes some webinars on English
Dataset structure:
mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music.
txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.ESpeech-webinars2
Webinar Audio Dataset
Dataset Description
This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Asessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.fleurs-regmix-webdataset
FLEURS RegMix WebDataset
Public, training-oriented WebDataset conversion of google/fleurs, pinned to source revision 70bb2e84b976b7e960aa89f1c648e09c59f894dd.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Shard names keep the source split, data/<language>/<language>-<split>-<index>.tar, so a training
mixture can be assembled without pulling the FLEURS evaluation splits into it. Every sample is a pair with the same key:
<key>.opus: mono Ogg… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/fleurs-regmix-webdataset.common-voice-17-regmix-webdataset
Common Voice 17 RegMix WebDataset
Public, training-oriented WebDataset conversion of fsicoli/common_voice_17_0, pinned to source revision 8262c16bf297c87a9cd88c51997c4758ed7a8ba2.
Layout
Each language is a RegMix cluster at data/<language>/*.tar. Every sample is a pair with the same key:
<key>.opus: mono Ogg Opus audio at 24 kbps, preserving the source sample rate (normally 48 kHz)
<key>.json: UTF-8 training metadata and the complete original TSV row
The source… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/common-voice-17-regmix-webdataset.asrtts_packed_webdataset
ASR+TTS Repacked Data (3.56M samples, mp3)
This dataset is a WebDataset repack prepared for FlexiSLM training (ASR+TTS tasks).
Paper: https://arxiv.org/abs/2606.31247
Demo page: https://flexislm.github.io/
Code: https://github.com/AmphionTeam/FlexiSLM
FlexiSLM-Data is a large-scale, single-turn English speech-to-speech dialogue dataset
for training FlexiSLM, a spoken language model.
This repository contains the paired prompt-and-response audio portion of the release in… See the full description on the dataset page: https://huggingface.co/datasets/FlexiSLM/asrtts_packed_webdataset.
