datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.open-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
openslr
madasr
shrutilipi
flerus
kathbath
indictts
ucla
gali
opendata-iisys-hui
HUI-Audio-Corpus-German Dataset
Overview
The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the Institute of Information Systems (IISYS). This dataset is designed to facilitate the development and training of TTS applications, particularly in the German language. The associated research paper can be found here.
Dataset Contents
The dataset comprises recordings from multiple speakers, with the five most… See the full description on the dataset page: https://huggingface.co/datasets/Paradoxia/opendata-iisys-hui.bn-asr-mega-open-dataopen-large-bengali-asr-data
Open Large Bengali ASR Data
This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).
Datasets:
commonvoice
open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.Kirundi_Open_Speech_Dataset
🇧🇮 Kirundi Open Speech & Text Dataset
Giving Kirundi a voice in AI.
📦 Dataset • 🫱🏿🫲🏾 Contribute • 📄 Files • 📚 Citation
📦 Dataset at a Glance
Public version (this repository)
Full dataset (private)
Content
Kirundi sentences still to translate, each with an AI draft (Machine_Suggestion)
Complete Kirundi ↔ French ↔ English pairs, audio and speaker metadata
Size
1,813 sentences
4,809 sentences, 2,996 complete pairs (62%)
Access
Free… See the full description on the dataset page: https://huggingface.co/datasets/Ijwi-ry-Ikirundi-AI/Kirundi_Open_Speech_Dataset.WanJuanSiLu-Multimodal-5Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-5Languages.open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Chinese Test Set (FLEURS)
This repository holds a test-only copy of the Mandarin Chinese test split of
google/fleurs (config cmn_hans_cn, split test).
It is used for the Chinese column of the Open ASR Leaderboard
(huggingface/open_asr_leaderboard#147).
The layout matches the FLEURS configs in
hf-audio/open-asr-leaderboard-multilingual-datasets,
so this config may later be merged into that repository.
How to load
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/open-asr-leaderboard-multilingual-datasets.open_data_asrcs_open_data_asropen-music-dataset-demo
Dataset Card for "open-music-dataset-demo"
More Information needed
open-bambara-asr-datasettendances-audio-video-barometre
Tendances audio-vidéo - Baromètre
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Tendances audio-vidéo - Baromètre qui est disponible à l'adresse https://www.data.gouv.fr/datasets/6836e0dc1baaf48fbb8b1851
Description
L’Arcom publie les données du volet quantitatif de son Baromètre Tendances audio-vidéo. L’étude est conduite auprès d’un échantillon représentatif de Français âgés de 15 ans et plus.
Elle vise à… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/tendances-audio-video-barometre.hallym_AI_OpenDataset
Hallym Adult and Child Speech Dataset
This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research.
Dataset Overview
Total Records: 2,714
Speakers: 49 (adult: 25, child: 24)
Groups: adult, child
File Format: WAV (audio) + TXT (transcription)
Speaker Statistics
Group
Count
Gender
Age Range
Adult
25명
남/여
50~78세
Child
24명
남/여
3~8세
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.WanJuanSiLu-Multimodal-3Languages
WanJuan·SiLu Multimodal Multilingual Corpus
🌏Dataset Introduction
The newly upgraded "Wanjuan·Silk Road Multimodal Corpus" brings the following three core improvements:
The number of languages has been significantly expanded: Based on the five open-source languages of "Wanjuan·Silk Road", namely Arabic, Russian, Korean, Vietnamese, and Thai, "Wanjuan·Silk Road Multimodal" has added three scarce corpus data of Serbian, Hungarian, and Czech, and uses the above eight key… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuanSiLu-Multimodal-3Languages.
