datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASRD-Dataset
Adversarial Surface-Form Robustness Dataset (ASRD)
💻 GitHub · 🤗 Dataset · 📦 Zenodo · 📝 Cite · 🛡️ Responsible Use
News
[2026/09] 🎉 Accepted at EvoRobust @ NeurIPS 2026, the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (Sydney, Australia).
[2026/10] 📦 v1.0.0 released and archived on Zenodo (10.5281/zenodo.23103902).
Paper: Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical… See the full description on the dataset page: https://huggingface.co/datasets/pavanmaddula/ASRD-Dataset.task963_librispeech_asr_next_word_prediction
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task963_librispeech_asr_next_word_prediction
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task963_librispeech_asr_next_word_prediction.gendl-hw1-asr-spell-ollama
Исправление ошибок распознавания речи — GenDL HW1
Сгенерированный мой датаест
Поля
Поле
Содержимое
input
Исходный текст, в котором могут быть ошибки распознавания.
output
Целевой исправленный текст.
Файлы
asr_spell_train.jsonl — основной датасет: одна JSON-запись на строку.
asr_spell_train.csv — копия тех же данных в CSV.
Загрузка
from datasets import load_dataset
dataset = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ImFastAsBlitz/gendl-hw1-asr-spell-ollama.task964_librispeech_asr_text_auto_completion
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task964_librispeech_asr_text_auto_completion
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task964_librispeech_asr_text_auto_completion.persian-asr-text-2.69M-deduped
🗂️ persian-asr-text-2.69M-deduped
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Deduplicated Persian ASR text dataset used by the training stack.
پیکرهٔ متنی فارسیِ حذفتکرارشده برای ساخت واژگان، مدلسازی زبانی و پشتیبانی از آموزش ASR.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
4 files; approximately 109.64 MB
4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.ko-finance-asr-corrections
ko-finance-asr-corrections
Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube.
210 pairs
mined from 2,391 videos of auto-captions
across 47 channels
totalling 1,080.1 hours
Each pair carries how often the term was mangled and how often it was said correctly, plus
verification provenance.
한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답
표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다.
What makes it different
No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.nb-asr-numerics-harvested
Norwegian Bokmål Numeric Expression Harvesting Dataset
This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn).
This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.genmodels-asr-corrections-hw1
Исправление ошибок распознавания речи
1351 пара input → output, полученная через Groq (openai/gpt-oss-120b). Правильные фразы — названия фильмов и имена; модель добавляет ошибки. В 270 примерах исправлять нечего.
Train: 1208 строк, validation: 143. Варианты одной правильной фразы находятся в одном разбиении.
nb-asr-numerics-categorized
Norwegian Bokmål Numeric Expression Categorized Dataset
This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1.
Source Dataset
Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards).
Processing Architecture
Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized.nb-asr-morfologisk-stil-nob
Document-conditioned Bokmål morphological style
This dataset trains a model to rewrite one sentence in the morphological style
described by 5–20 observed word forms from the same source document. Every
listed form has verified alternatives in AltMorph. Evidence is deduplicated,
comes from another sentence, and is disjoint from the target's alternative
families.
Task format
The prompt is ready for sequence-to-sequence training:
skriv_i_samme_morfologiske_stil:… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-morfologisk-stil-nob.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.uzbek-customer-support-dialogs
Uzbek Customer Support Dialogs 🇺🇿
A high-quality dataset of 990 customer support conversations in Uzbek (Latin script), designed for training and fine-tuning conversational AI models.
This is one of the first large-scale customer support datasets in Uzbek, created to address the gap of low-resource NLP for Central Asian languages.
📋 Dataset Description
990 conversational dialogs in natural Uzbek (Latin script)
11 customer support categories: Order, Shipping, Cancel… See the full description on the dataset page: https://huggingface.co/datasets/AsrorAsr/uzbek-customer-support-dialogs.cleaned-asr-transcriptsaerograph-asrs
AeroGraph ASRS Dataset
2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports
with LLM-extracted entities and relations for knowledge graph construction.
Dataset Description
This dataset contains processed ASRS incident narratives along with
structured entity and relation extractions conforming to an aviation
safety ontology (10 entity types, 8 edge types).
Reports Split
2000 reports from the NASA ASRS database
Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.mondegreen-asr-errors
Mondegreen ASR error pairs
(ASR hypothesis, gold text) pairs for Japanese ASR post-correction.
This build is simulated -- errors come from a phonetic corruption model, not from a real ASR system. It exists so the whole pipeline (gate training, benchmarks, figures, CI) is reproducible without a GPU. Treat every number derived from it as a stated assumption, not a measurement.
How it was made
synthetic text
-> phonetic corruption model (mondegreen.simulate)
->… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/mondegreen-asr-errors.nb-asr-numerics-categorized-smoke-test
Norwegian Bokmål Numeric Expression Categorized Dataset
This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1.
Source Dataset
Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards).
Processing Architecture
Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.nb-asr-numerics-balanced
Balanced Synthetic Norwegian Bokmål Numerics Dataset
This dataset provides a class-balanced synthetic corpus of Norwegian Bokmål sentences containing numeric expressions. It draws 10,000 examples for each of the 59 numeric categories (totaling 590,000 rows).
Source & Synthesis Architecture
Templates source: pere/nb-asr-numerics-categorized.
Methodology:
Filtered the original dataset for kept rows containing annotated entities.
For each target category, sampled 10… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-balanced.ASRS-ChatGPT
Dataset Summary
The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives.
The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available.
Dataset Structure
The column names present in the dataset and their descriptions are provided below:
Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.task965_librispeech_asr_missing_word_prediction
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task965_librispeech_asr_missing_word_prediction
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task965_librispeech_asr_missing_word_prediction.fluid-2-sft-asr
Fluid 2 — synthetic dictation cleanup
Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB).
This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.kikongo-bible-asr-embeddings
Kikongo Bible Embeddings
This dataset is a version of the kikongo-bible-asr dataset. I used the cohere-emdbed-v3 model to produce the embeddings.
cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.ASR-REF-PRED
