datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hindi-asr-wdsTaiwan-Tongues-ASR-CE-dataset-hokkien
Taiwan-Tongues-ASR-CE-dataset-hokkien
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hokkien.Galgame_Speech_ASR_16kHz
Dataset Card for Galgame_Speech_ASR_16kHz
[!IMPORTANT]The following rules (in the original repository) must be followed:
必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关!
训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。
English:
You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_ASR_16kHz.WSYue-ASR-eval
WSYue-ASR-eval: Cantonese ASR Benchmark
To address the unique linguistic characteristics of Cantonese in speech recognition, we propose WSYue-ASR-eval, a benchmark specifically designed for evaluating Cantonese ASR systems. It is tailored to assess model performance across diverse lengths, domains, and linguistic phenomena of Cantonese speech.
The test set annotations are provided by Beijing AISHELL Technology Co., Ltd.
Key features:
Annotated through multiple rounds of manual… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WSYue-ASR-eval.Taiwan-Tongues-ASR-CE-dataset-zhtw
Taiwan-Tongues-ASR-CE-dataset-zhtw
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-zhtw.ganjoor-chunked-asr-datasetvocal-burst-annotation-asr-tuning-dataset
Vocal Burst Annotation ASR Tuning Dataset
A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information.
Example Transcript
[nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.rohingya_asr_audioThis is the first public Rohingya language ASR dataset in AI history.
Overview
This dataset contains broadcast audio recordings from the Voice of America (VOA) Rohingya Service. Each file represents a daily news segment, typically 30 minutes in length, automatically segmented into chunks of 5–15 seconds for use in self-supervised ASR, pretraining, language identification, and more.
The content was aired publicly as part of VOA’s Rohingya-language radio program and is therefore… See the full description on the dataset page: https://huggingface.co/datasets/freococo/rohingya_asr_audio.ganjoor-datasetvoa_myanmar_asr_audio_1
📢 This is the first publicly released ASR-ready Burmese speech dataset with over 1 million audio chunks — a milestone in the history of Myanmar language technology.
Overview
This dataset was created by scraping and segmenting the full archive of the VOA Burmese morning radio program. Out of a total of 3,687 full-length MP3 broadcasts, this release processes 3,267 of them, resulting in approximately 1.8 million sentence-level audio chunks, totaling ~3,267 hours of segmented audio.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_1.Taiwan-Tongues-ASR-CE-dataset-hakka
Taiwan-Tongues-ASR-CE-dataset-hakka
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-hakka.115hours_pvtv_myanmar_asr
115 Hours PVTV Myanmar ASR
This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps.
Dedication
This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.nug_myanmar_asr
366 Hours NUG Myanmar ASR Dataset
The NUG Myanmar ASR Dataset is the first large-scale open Burmese speech dataset — now expanded to over 521,476 audio-text pairs, totaling ~366 hours of clean, segmented audio. All data was collected from public-service educational broadcasts by the National Unity Government (NUG) of Myanmar and the FOEIM Academy.
This dataset is released under a CC0 1.0 Universal license — fully open and public domain. No attribution required.
🕊️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/nug_myanmar_asr.shan_language_asr_voices
⭐ A Voice for the Shan People: The SHAN Herald Agency Audio Archive
This is an extensive 306-hour audio dataset of the Shan (Tai-Yai) language, meticulously curated from the public broadcasts of the Shan Herald Agency for News (SHAN). For over two decades, SHAN has been a vital, independent voice for the people of Shan State, Myanmar, chronicling their stories of culture, politics, and the enduring struggle for federal democracy.
This archive stands as one of the largest publicly… See the full description on the dataset page: https://huggingface.co/datasets/freococo/shan_language_asr_voices.Taiwan-Tongues-ASR-CE-dataset-en
Taiwan-Tongues-ASR-CE-dataset-en
本資料集為 Taiwan-Tongues-ASR-CE 專案所使用的預訓練資料,透過 WebDataset 格式打包,並上傳至 Hugging Face 以便研究人員與開發者自由取用。
📂 Dataset 結構
本資料集分為 Training 與 Test 兩個子集,均以 WebDataset tar 檔案形式存放:
Training set (WebDataset format)
train/train-000000.tar
train/train-000001.tar
...
Test set (WebDataset format)
test/test-000000.tar
...
tsv set
train.tsv
test.tsv
...
每個 tar 內部均包含對應的音檔與標註,方便直接搭配 WebDataset 與 PyTorch / Hugging Face datasets 進行訓練與測試。
🏷️… See the full description on the dataset page: https://huggingface.co/datasets/adi-gov-tw/Taiwan-Tongues-ASR-CE-dataset-en.noisy-join-mixed-asr
Noisy join mixed ASR
Problem with common ASR dataset, each audio sample is a monolanguage, but we can generate an audio sample with multilanguage using ASR and Force Alignment models. Notebooks at https://github.com/huseinzol05/malaya-speech/tree/master/data/noisy-join-mixed-asr
Figure like below,
Generated samples ~601 hours.
ASR_Fellowship_Challenge_Datasetlibrispeech_asr_for_speaker_turnsagaw_karen_asrThis is the first public Sagaw Karen language ASR dataset in AI history.
Sagaw Karen ASR
This dataset contains audio recordings and aligned metadata in the Sagaw Karen language (ISO 639-3: ksw), a major Sgaw Karenic language spoken throughout southern and eastern Myanmar. The language is sometimes also referred to as Sgaw Karen or Sakaw Karen in English transliterations.
All audio segments in this dataset were sourced from publicly available news broadcasts published by PVTV… See the full description on the dataset page: https://huggingface.co/datasets/freococo/sagaw_karen_asr.zomi_asrThis is the first public Zomi language ASR dataset in AI history.
Zomi ASR
This dataset contains audio recordings and aligned metadata in the Zomi language — a collective ethnolinguistic identity adopted by some Kuki-Chin language-speaking communities in Myanmar and India. The term Zomi means "Zo people", derived from the root word Zo (ancestral identity) and mi meaning "people." While originally coined to encompass all Zo-related communities, usage of the term varies regionally and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/zomi_asr.mon_language_asr_audio
RFA Mon Language Voices
This dataset contains 14.8 hours of audio in the Mon language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Mon language, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented into 3,634 manageable chunks and… See the full description on the dataset page: https://huggingface.co/datasets/freococo/mon_language_asr_audio.mozilla-common-voice-spontaneous-speech-asr-shared-task
Mozilla Common Voice Spontaneous Speech ASR Shared Task
This repository combines the Mozilla Data Collective Common Voice spontaneous speech ASR shared-task
train/dev and test archives in one place.
Locales present across the combined train/dev and test packages: ady, aln, bas, bew, bxk,
cgg, el-CY, hch, kbd, kcn, koo, led, lke, lth, meh, mmc, pne, qxp, ruc,
rwm, sco, tob, top, ttj, ukv, ush.
Split package
Mozilla Data Collective dataset ID
Hub archive
Original MDC archive… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/mozilla-common-voice-spontaneous-speech-asr-shared-task.samromur_asr
Dataset Card for samromur_asr
Dataset Summary
This is a modfied copy of the dataset from The Language and Voice Laboratory in RU.
This is the first release of the Samrómur Icelandic Speech corpus that contains 100.000 validated utterances.
The corpus is a result of the crowd-sourcing effort run by the Language and Voice Lab at the Reykjavik University, in cooperation with Almannarómur, Center for Language Technology.
Languages
The audio is in Icelandic.
The… See the full description on the dataset page: https://huggingface.co/datasets/DavidErikMollberg/samromur_asr.western_poe_karen_asrThis is the first public Western Poe Karen language ASR dataset in AI history.
Western Poe Karen ASR
This dataset contains audio recordings and aligned transcriptions in the Western Poe Karen language (also known in linguistic literature as Western Pwo or Delta Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in the Ayeyarwady Delta region of Myanmar. Although linguists commonly refer to this language as Western Pwo Karen, the community and this project prefer the spelling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/western_poe_karen_asr.karenni_language_asr_audio
RFA Karenni (Kayah) Language Voices
This dataset contains 17 hours of audio in the Karenni (Kayah) language, sourced from news broadcasts by Radio Free Asia (RFA) Burmese. This is one of the largest publicly accessible audio resources for the Karenni language family, designed to support research in low-resource automatic speech recognition (ASR), voice activity detection, and other speech-related tasks.
This dataset was created by freococo.
The audio has been automatically segmented… See the full description on the dataset page: https://huggingface.co/datasets/freococo/karenni_language_asr_audio.kachin_asr_audio
Dataset Summary
This is the first public Kachin language ASR dataset in history.
Kachin ASR Audio is a collection of speech data in the Kachin (Jinghpaw) language, sourced entirely from publicly available PVTV (People’s Voice Television) broadcasts. The dataset includes narration, interviews, and spoken reports intended to support the development of automatic speech recognition (ASR) systems for rare-resource indigenous languages in Myanmar.
Each audio file is paired with metadata… See the full description on the dataset page: https://huggingface.co/datasets/freococo/kachin_asr_audio.google_myanmar_asr_voices
Google Myanmar ASR Dataset (WebDataset Version)
This repository provides a clean, user-friendly, and robust version of the Google Myanmar ASR Dataset, which is derived from the OpenSLR-80 Burmese Speech Corpus.
This version has been carefully re-processed into the WebDataset format. Each sample consists of a .wav audio file and a clean .json metadata file, packaged into sharded .tar archives. This format is highly efficient for large-scale training of ASR models.… See the full description on the dataset page: https://huggingface.co/datasets/freococo/google_myanmar_asr_voices.eastern_poe_karen_asrThis is the first public Eastern Poe Karen language ASR dataset in AI history.
Eastern Poe Karen ASR
This dataset contains audio recordings and aligned metadata in the Eastern Poe Karen language (a regional variety of Eastern Pwo, ISO 639-3: pwo), a Karenic language spoken primarily in Mon State and Kayin State in southeastern Myanmar. While linguistically described as Eastern Pwo Karen, the community and this project prefer the term Poe as a community-endorsed spelling.
All audio… See the full description on the dataset page: https://huggingface.co/datasets/freococo/eastern_poe_karen_asr.voa_myanmar_asr_audio_2⸻
Overview
This dataset was created by scraping and segmenting over 4,000 episodes of the VOA Burmese morning radio program. From that archive, 3,687 MP3 files were extracted and processed. This dataset contains sentence-level audio chunks suitable for ASR and speech-related model training.
The current release (voa_batch_001.tar and voa_batch_003.tar) contains a combined total of ~152,300 sentence-level audio chunks derived from the first 420 MP3 files in the archive, totaling… See the full description on the dataset page: https://huggingface.co/datasets/freococo/voa_myanmar_asr_audio_2.NEPALI-ASR
