datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.asr-leaderboard-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.ASR-datasets-ptbr
📚 Datasets de Áudio em Português (PT-BR)
Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition).
O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade.
📂 Datasets Integrados
A tabela abaixo lista todos os datasets incluídos, com suas informações:
Dataset
Config Name
TOTAL
train
test
validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.Datasets_ENsea_audiobench_datasets_ASR
SEA-SpeechBench — ASR (Automatic Speech Recognition)
This dataset is the automatic-speech-recognition (ASR) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 26,863 evaluation examples
across eleven languages, drawn from fifteen source corpora, each pairing an
audio recording with an instruction and a reference transcript.
Given the recording and the instruction, a model must transcribe the speech.… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_ASR.sea_audiobench_datasetssea_audiobench_datasets_TCQ
SEA-SpeechBench — TCQ (Timestamped Content Query)
This dataset is the timestamped-content-query (TCQ) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 14,172 evaluation examples
across five languages, each pairing an audio recording with an instruction
and a reference answer.
Given the recording and a timestamp, a model must report what is said at
that point in the audio. Contexts run from 30 seconds to 3… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TCQ.doc-audio-6
[doc] audio dataset 6
This dataset contains 4 audio files in the /train directory, with a CSV metadata file providing another data column.
Punjabi_ASR_datasetsOpenS2S_Datasets
How to Use?
Download, merge the files, and extract
You can run the following command to merge the compressed file parts after downloading.
cat en_response_wav.tar.gz.* > en_response_wav.tar.gz
cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz
IMDA-NSC-datasetssea_audiobench_datasets_PQA
SEA-SpeechBench — Paralinguistics (AGE, ER, GR)
This dataset holds the three paralinguistic tasks of SEA-SpeechBench, a
large-scale multitask benchmark for speech understanding across Southeast
Asia. It contains 25,563 evaluation examples across nine languages, drawn from
fourteen source corpora, each pairing an audio recording with an instruction
and a reference answer.
Given the recording and the instruction, a model must identify a property of
the speaker or the delivery… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_PQA.rhapsodie
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/corpora/rhapsodie
Description
Corpus de français parlé annoté pour la prosodie et la syntaxe
Un problème central dans l’étude des langues parlées est la compréhension du rôle que jouent les indices intonosyntaxiques dans la segmentation du continuum sonore en unités informationnelles et discursives. Se posent notamment les questions suivantes : quel est le degré de congruence entre les différentes unités manipulées par la… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/rhapsodie.audiobench_datasetsmulti-stream-spontaneous-conversation-training-datasets_chinese
Multi-stream Spontaneous Conversation Training Datasets_Chinese
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: Chinese
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_chinese.multi-stream-spontaneous-conversation-training-datasets_chinese
Multi-stream Spontaneous Conversation Training Datasets_Chinese
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: Chinese
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/multi-stream-spontaneous-conversation-training-datasets_chinese.audiofolder_two_configs_in_metadataaudiofolder_single_config_in_metadatapreprocessed_speech_datasetssea_audiobench_datasets_TLoc
SEA-SpeechBench — TLoc (Temporal Localization)
This dataset is the temporal-localization (TLoc) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 6,736 evaluation examples
across five languages, each pairing an audio recording of 30–180 seconds
with an instruction and a reference answer.
Given the recording and the instruction, a model must identify when a
described utterance occurs.
Quick start… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TLoc.InsectSet459
InsectSet459: An Open Dataset of Insect Sounds for Bioacoustic Machine Learning
Overview
InsectSet459 is a comprehensive dataset of insect sounds designed for developing and testing machine learning algorithms for automatic insect identification. It contains 26,399 audio files from 459 species of Orthoptera (crickets, grasshoppers, katydids) and Cicadidae (cicadas), providing 9.5 days of audio material.
Key Features
459 unique insect species (310 Orthopteran… See the full description on the dataset page: https://huggingface.co/datasets/academic-datasets/InsectSet459.asr-leaderboard-datasets
Afrivoice and Amharic configs
These configs were added by _data-prep-gsma/prep_final.py and
_data-prep-gsma/prep_amharic.py in the gsma-asr-bench
project. All Afrivoice rows are test-only: each config below contains
exactly the held-out evaluation partition from the upstream dataset.
Common schema (identical across the three configs):
column
type
notes
file_name
string
stable per-config identifier
audio
Audio(sampling_rate=16000)
mono
duration
float64
seconds… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-leaderboard-datasets.multi-stream-spontaneous-conversation-training-datasets_english
Multi-stream Spontaneous Conversation Training Datasets_English
Every data point counts.
Dataset Basic Info
Dataset Type: ASR Corpus
Language: English
Audio Parameters: 16 kHz, 16 bits
File Format: WAV (PCM)
Recording Equipment: Mobile device
Dataset Description
The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_english.sea_audiobench_datasets_SQA
SEA-SpeechBench — SQA (Spoken Question Answering)
This dataset is the spoken-question-answering (SQA) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 5,462 evaluation examples
across five languages, each pairing an audio recording with a question and a
reference answer.
Given the recording and the question, a model must answer using the content
of the speech.
Quick start
Requires datasets>=4.0… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_SQA.audiofolder_no_configs_in_metadatasea_audiobench_datasets_ST
SEA-SpeechBench — ST (Speech Translation)
This dataset is the speech-translation (ST) task of SEA-SpeechBench, a
large-scale multitask benchmark for speech understanding across Southeast
Asia. It contains 7,189 evaluation examples across nine source languages,
each pairing an audio recording with an instruction and a reference
translation.
Given the recording and the instruction, a model must translate the speech
into the target language.
Quick start
Requires… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_ST.Rogue-Datasets
Rogue RVC Voice Dataset
Audio dataset prepared for training a Retrieval-Based Voice Conversion (RVC) model of Rogue with Applio.
Repository:
https://huggingface.co/datasets/0xra/Rogue-Datasets
This is a voice-conversion training dataset, not a text-to-speech corpus. The supplied archive contains segmented WAV clips and training metadata; it does not contain transcriptions.
Dataset summary
Property
Value
Audio clips
626
Total duration
35:00.422… See the full description on the dataset page: https://huggingface.co/datasets/0xra/Rogue-Datasets.speaker_datasets_parler_v1open-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.
