Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01D4nt3 /esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset. def add_duration(sample): y, sr = sample['audio']["array"], sample['audio']["sampling_rate"] sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000 return sample tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True) # compute duration to filter tedlium = tedlium.map(add_duration) tedlium = tedlium.select(range(512)) # Whisper max supported duration tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.audion<1K0 likes9.9k downloads2y agoHugging Face02nithinraok /asr-leaderboard-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.audioautomatic-speech-recognition100K<n<1M4 likes5.4k downloads1y agoHugging Face03opedromartins /ASR-datasets-ptbr 📚 Datasets de Áudio em Português (PT-BR) Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition). O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade. 📂 Datasets Integrados A tabela abaixo lista todos os datasets incluídos, com suas informações: Dataset Config Name TOTAL train test validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.audioautomatic-speech-recognition1M<n<10M13 likes2.3k downloads1y agoHugging Face04hf-audio /open-asr-leaderboard-multilingual-datasets ASR Leaderboard Datasets This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS). How to Load To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>. from datasets import load_dataset # Load the FLEURS dataset for Bulgarian fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg") print(fleurs_bg) # Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition100K<n<1M4 likes1.5k downloads3mo agoHugging Face05tanooki426 /Datasets_ENaudioautomatic-speech-recognition1K<n<10K4 likes828 downloads11d agoHugging Face06MERaLiON /sea_audiobench_datasets_ASR SEA-SpeechBench — ASR (Automatic Speech Recognition) This dataset is the automatic-speech-recognition (ASR) task of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 26,863 evaluation examples across eleven languages, drawn from fifteen source corpora, each pairing an audio recording with an instruction and a reference transcript. Given the recording and the instruction, a model must transcribe the speech.… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_ASR.audioautomatic-speech-recognition10K<n<100K0 likes723 downloads1mo agoHugging Face07zxl /sea_audiobench_datasetsaudio100K<n<1M0 likes668 downloads1y agoHugging Face08MERaLiON /sea_audiobench_datasets_TCQ SEA-SpeechBench — TCQ (Timestamped Content Query) This dataset is the timestamped-content-query (TCQ) task of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 14,172 evaluation examples across five languages, each pairing an audio recording with an instruction and a reference answer. Given the recording and a timestamp, a model must report what is said at that point in the audio. Contexts run from 30 seconds to 3… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TCQ.audioquestion-answering10K<n<100K0 likes623 downloads1mo agoHugging Face09datasets-examples /doc-audio-6 [doc] audio dataset 6 This dataset contains 4 audio files in the /train directory, with a CSV metadata file providing another data column. audion<1K0 likes488 downloads2y agoHugging Face10kdcyberdude /Punjabi_ASR_datasetsaudio100K<n<1M7 likes454 downloads2y agoHugging Face11CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes449 downloads1y agoHugging Face12chanchungkit /IMDA-NSC-datasetsaudio1M<n<10M0 likes417 downloads1y agoHugging Face13MERaLiON /sea_audiobench_datasets_PQA SEA-SpeechBench — Paralinguistics (AGE, ER, GR) This dataset holds the three paralinguistic tasks of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 25,563 evaluation examples across nine languages, drawn from fourteen source corpora, each pairing an audio recording with an instruction and a reference answer. Given the recording and the instruction, a model must identify a property of the speaker or the delivery… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_PQA.audioaudio-classification10K<n<100K0 likes413 downloads1mo agoHugging Face14datasets-CNRS /rhapsodie [!NOTE] Dataset origin: https://www.ortolang.fr/market/corpora/rhapsodie Description Corpus de français parlé annoté pour la prosodie et la syntaxe Un problème central dans l’étude des langues parlées est la compréhension du rôle que jouent les indices intonosyntaxiques dans la segmentation du continuum sonore en unités informationnelles et discursives. Se posent notamment les questions suivantes : quel est le degré de congruence entre les différentes unités manipulées par la… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/rhapsodie.audion<1K0 likes405 downloads2y agoHugging Face15zxl /audiobench_datasetsaudio100K<n<1M1 likes378 downloads1y agoHugging Face16MagicHub /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes359 downloads4mo agoHugging Face17MagicDataTech /multi-stream-spontaneous-conversation-training-datasets_chinese Multi-stream Spontaneous Conversation Training Datasets_Chinese Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: Chinese Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicDataTech/multi-stream-spontaneous-conversation-training-datasets_chinese.audio1K<n<10K2 likes334 downloads4mo agoHugging Face18datasets-maintainers /audiofolder_two_configs_in_metadataaudion<1K0 likes321 downloads3y agoHugging Face19datasets-maintainers /audiofolder_single_config_in_metadataaudion<1K0 likes302 downloads3y agoHugging Face20qaz159qaz159 /preprocessed_speech_datasetsaudio100K<n<1M0 likes284 downloads2y agoHugging Face21MERaLiON /sea_audiobench_datasets_TLoc SEA-SpeechBench — TLoc (Temporal Localization) This dataset is the temporal-localization (TLoc) task of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 6,736 evaluation examples across five languages, each pairing an audio recording of 30–180 seconds with an instruction and a reference answer. Given the recording and the instruction, a model must identify when a described utterance occurs. Quick start… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TLoc.audioquestion-answering1K<n<10K0 likes281 downloads1mo agoHugging Face22academic-datasets /InsectSet459 InsectSet459: An Open Dataset of Insect Sounds for Bioacoustic Machine Learning Overview InsectSet459 is a comprehensive dataset of insect sounds designed for developing and testing machine learning algorithms for automatic insect identification. It contains 26,399 audio files from 459 species of Orthoptera (crickets, grasshoppers, katydids) and Cicadidae (cicadas), providing 9.5 days of audio material. Key Features 459 unique insect species (310 Orthopteran… See the full description on the dataset page: https://huggingface.co/datasets/academic-datasets/InsectSet459.audioaudio-classification10K<n<100K2 likes253 downloads2y agoHugging Face23SaarAI /asr-leaderboard-datasetsgated Afrivoice and Amharic configs These configs were added by _data-prep-gsma/prep_final.py and _data-prep-gsma/prep_amharic.py in the gsma-asr-bench project. All Afrivoice rows are test-only: each config below contains exactly the held-out evaluation partition from the upstream dataset. Common schema (identical across the three configs): column type notes file_name string stable per-config identifier audio Audio(sampling_rate=16000) mono duration float64 seconds… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-leaderboard-datasets.audio100K<n<1M0 likes248 downloads12d agoHugging Face24MagicHub /multi-stream-spontaneous-conversation-training-datasets_english Multi-stream Spontaneous Conversation Training Datasets_English Every data point counts. Dataset Basic Info Dataset Type: ASR Corpus Language: English Audio Parameters: 16 kHz, 16 bits File Format: WAV (PCM) Recording Equipment: Mobile device Dataset Description The Multi-stream conversation dataset developed by MagicData captures each speaker's audio track and labels each speaker separately, thereby preserving the natural occurrences of… See the full description on the dataset page: https://huggingface.co/datasets/MagicHub/multi-stream-spontaneous-conversation-training-datasets_english.audio1K<n<10K1 likes203 downloads4mo agoHugging Face25MERaLiON /sea_audiobench_datasets_SQA SEA-SpeechBench — SQA (Spoken Question Answering) This dataset is the spoken-question-answering (SQA) task of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 5,462 evaluation examples across five languages, each pairing an audio recording with a question and a reference answer. Given the recording and the question, a model must answer using the content of the speech. Quick start Requires datasets>=4.0… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_SQA.audioquestion-answering1K<n<10K0 likes145 downloads1mo agoHugging Face26datasets-maintainers /audiofolder_no_configs_in_metadataaudion<1K0 likes124 downloads3y agoHugging Face27MERaLiON /sea_audiobench_datasets_ST SEA-SpeechBench — ST (Speech Translation) This dataset is the speech-translation (ST) task of SEA-SpeechBench, a large-scale multitask benchmark for speech understanding across Southeast Asia. It contains 7,189 evaluation examples across nine source languages, each pairing an audio recording with an instruction and a reference translation. Given the recording and the instruction, a model must translate the speech into the target language. Quick start Requires… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_ST.audiotranslation1K<n<10K0 likes122 downloads1mo agoHugging Face280xra /Rogue-Datasets Rogue RVC Voice Dataset Audio dataset prepared for training a Retrieval-Based Voice Conversion (RVC) model of Rogue with Applio. Repository: https://huggingface.co/datasets/0xra/Rogue-Datasets This is a voice-conversion training dataset, not a text-to-speech corpus. The supplied archive contains segmented WAV clips and training metadata; it does not contain transcriptions. Dataset summary Property Value Audio clips 626 Total duration 35:00.422… See the full description on the dataset page: https://huggingface.co/datasets/0xra/Rogue-Datasets.audioaudio-to-audion<1K0 likes120 downloads14d agoHugging Face29ylacombe /speaker_datasets_parler_v1audio10K<n<100K0 likes116 downloads2y agoHugging Face30Metric-AI /open-asr-leaderboard-multilingual-datasets Open ASR Leaderboard Armenian Test Datasets This private repository holds leaderboard-compatible Armenian test configurations while their integration is being validated. Configurations fleurs_hy Source: google/fleurs, configuration hy_am, test split Reviewed reference changes: Metric-AI/fleurs-corrections, test split 932 recordings; all 314 reviewed corrections were matched to the original source transcript and applied mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.audioautomatic-speech-recognition1K<n<10K1 likes115 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.