Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pavanmaddula /ASRD-Datasetgated Adversarial Surface-Form Robustness Dataset (ASRD) 💻 GitHub · 🤗 Dataset · 📦 Zenodo · 📝 Cite · 🛡️ Responsible Use News [2026/09] 🎉 Accepted at EvoRobust @ NeurIPS 2026, the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (Sydney, Australia). [2026/10] 📦 v1.0.0 released and archived on Zenodo (10.5281/zenodo.23103902). Paper: Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical… See the full description on the dataset page: https://huggingface.co/datasets/pavanmaddula/ASRD-Dataset.texttext-generation1K<n<10K0 likes183 downloads8d agoHugging Face02Lots-of-LoRAs /task963_librispeech_asr_next_word_prediction Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task963_librispeech_asr_next_word_prediction Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task963_librispeech_asr_next_word_prediction.texttext-generationn<1K0 likes81 downloads2y agoHugging Face03ImFastAsBlitz /gendl-hw1-asr-spell-ollama Исправление ошибок распознавания речи — GenDL HW1 Сгенерированный мой датаест Поля Поле Содержимое input Исходный текст, в котором могут быть ошибки распознавания. output Целевой исправленный текст. Файлы asr_spell_train.jsonl — основной датасет: одна JSON-запись на строку. asr_spell_train.csv — копия тех же данных в CSV. Загрузка from datasets import load_dataset dataset = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ImFastAsBlitz/gendl-hw1-asr-spell-ollama.texttext-generation1K<n<10K0 likes76 downloads14d agoHugging Face04Lots-of-LoRAs /task964_librispeech_asr_text_auto_completion Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task964_librispeech_asr_text_auto_completion Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task964_librispeech_asr_text_auto_completion.texttext-generationn<1K0 likes71 downloads2y agoHugging Face05Reza2kn /persian-asr-text-2.69M-deduped 🗂️ persian-asr-text-2.69M-deduped English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Deduplicated Persian ASR text dataset used by the training stack. پیکرهٔ متنی فارسیِ حذف‌تکرارشده برای ساخت واژگان، مدل‌سازی زبانی و پشتیبانی از آموزش ASR. 🧩 Role Persian text and linguistic asset مصنوع متنی و زبانی فارسی 📦 Snapshot 4 files; approximately 109.64 MB 4 فایل؛ حدود 109.64… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-text-2.69M-deduped.tabularautomatic-speech-recognition1M<n<10M0 likes65 downloads2mo agoHugging Face06woongstar /ko-finance-asr-corrections ko-finance-asr-corrections Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube. 210 pairs mined from 2,391 videos of auto-captions across 47 channels totalling 1,080.1 hours Each pair carries how often the term was mangled and how often it was said correctly, plus verification provenance. 한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답 표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다. What makes it different No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.tabulartext-generationn<1K0 likes62 downloads29d agoHugging Face07pere /nb-asr-numerics-harvested Norwegian Bokmål Numeric Expression Harvesting Dataset This dataset contains cleaned, high-quality Norwegian Bokmål sentences containing numeric expressions harvested from both the Norwegian Colossal Corpus (NbAiLab/NCC) and the FineWeb-2 Norwegian Bokmål subset (HuggingFaceFW/fineweb-2 config nob_Latn). This is a combined high-volume intermediate dataset built for the first stage of a template-based Norwegian synthetic speech (TTS) generation pipeline to improve number/digit… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-harvested.tabulartext-generation1M<n<10M0 likes52 downloads3mo agoHugging Face08ASR2005Bluesnow /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.texttext-classification100K<n<1M0 likes49 downloads1mo agoHugging Face09Kazalika /genmodels-asr-corrections-hw1 Исправление ошибок распознавания речи 1351 пара input → output, полученная через Groq (openai/gpt-oss-120b). Правильные фразы — названия фильмов и имена; модель добавляет ошибки. В 270 примерах исправлять нечего. Train: 1208 строк, validation: 143. Варианты одной правильной фразы находятся в одном разбиении. texttext-generation1K<n<10K0 likes42 downloads7d agoHugging Face10pere /nb-asr-numerics-categorized Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized.tabulartext-generation1M<n<10M1 likes39 downloads3mo agoHugging Face11NbAiLab /nb-asr-morfologisk-stil-nob Document-conditioned Bokmål morphological style This dataset trains a model to rewrite one sentence in the morphological style described by 5–20 observed word forms from the same source document. Every listed form has verified alternatives in AltMorph. Evidence is deduplicated, comes from another sentence, and is disjoint from the target's alternative families. Task format The prompt is ready for sequence-to-sequence training: skriv_i_samme_morfologiske_stil:… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-morfologisk-stil-nob.text-generation0 likes36 downloads1mo agoHugging Face12bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes32 downloads5mo agoHugging Face13AsrorAsr /uzbek-customer-support-dialogs Uzbek Customer Support Dialogs 🇺🇿 A high-quality dataset of 990 customer support conversations in Uzbek (Latin script), designed for training and fine-tuning conversational AI models. This is one of the first large-scale customer support datasets in Uzbek, created to address the gap of low-resource NLP for Central Asian languages. 📋 Dataset Description 990 conversational dialogs in natural Uzbek (Latin script) 11 customer support categories: Order, Shipping, Cancel… See the full description on the dataset page: https://huggingface.co/datasets/AsrorAsr/uzbek-customer-support-dialogs.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face14bingbangboom /cleaned-asr-transcriptstexttext-generation10K<n<100K1 likes29 downloads7mo agoHugging Face15Aryan95614 /aerograph-asrs AeroGraph ASRS Dataset 2,000 real NASA Aviation Safety Reporting System (ASRS) incident reports with LLM-extracted entities and relations for knowledge graph construction. Dataset Description This dataset contains processed ASRS incident narratives along with structured entity and relation extractions conforming to an aviation safety ontology (10 entity types, 8 edge types). Reports Split 2000 reports from the NASA ASRS database Fields: id, text, aircraft_type… See the full description on the dataset page: https://huggingface.co/datasets/Aryan95614/aerograph-asrs.tabularquestion-answering1K<n<10K1 likes24 downloads6mo agoHugging Face16NagaYu /mondegreen-asr-errors Mondegreen ASR error pairs (ASR hypothesis, gold text) pairs for Japanese ASR post-correction. This build is simulated -- errors come from a phonetic corruption model, not from a real ASR system. It exists so the whole pipeline (gate training, benchmarks, figures, CI) is reproducible without a GPU. Treat every number derived from it as a stated assumption, not a measurement. How it was made synthetic text -> phonetic corruption model (mondegreen.simulate) ->… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/mondegreen-asr-errors.automatic-speech-recognition1K<n<10K0 likes21 downloads2mo agoHugging Face17pere /nb-asr-numerics-categorized-smoke-test Norwegian Bokmål Numeric Expression Categorized Dataset This dataset represents Stage 2 of the Norwegian numerics-data pipeline. It contains semantic validation and categorization annotations of the Norwegian numeric expression sentences harvested in Stage 1. Source Dataset Harvested Dataset: pere/nb-asr-numerics-harvested (approx. 3.8 million rows across 6 shards). Processing Architecture Inference Model: google/gemma-4-12B-it (instruction-tuned… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-categorized-smoke-test.tabulartext-generationn<1K0 likes18 downloads3mo agoHugging Face18pere /nb-asr-numerics-balanced Balanced Synthetic Norwegian Bokmål Numerics Dataset This dataset provides a class-balanced synthetic corpus of Norwegian Bokmål sentences containing numeric expressions. It draws 10,000 examples for each of the 59 numeric categories (totaling 590,000 rows). Source & Synthesis Architecture Templates source: pere/nb-asr-numerics-categorized. Methodology: Filtered the original dataset for kept rows containing annotated entities. For each target category, sampled 10… See the full description on the dataset page: https://huggingface.co/datasets/pere/nb-asr-numerics-balanced.texttext-generation100K<n<1M0 likes16 downloads3mo agoHugging Face19archanatikayatray /ASRS-ChatGPTgated Dataset Summary The dataset contains a total of 9984 incident records and 9 columns. Some of the columns contain ground truth values whereas others contain information generated by ChatGPT based on the incident Narratives. The creation of this dataset is aimed at providing researchers with columns generated by using ChatGPT API which is not freely available. Dataset Structure The column names present in the dataset and their descriptions are provided below: Column… See the full description on the dataset page: https://huggingface.co/datasets/archanatikayatray/ASRS-ChatGPT.textzero-shot-classification1K<n<10K8 likes14 downloads3y agoHugging Face20Lots-of-LoRAs /task965_librispeech_asr_missing_word_prediction Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task965_librispeech_asr_missing_word_prediction Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task965_librispeech_asr_missing_word_prediction.texttext-generationn<1K0 likes11 downloads2y agoHugging Face21johnbean393 /fluid-2-sft-asrgated Fluid 2 — synthetic dictation cleanup Fluid 2 is an English supervised-fine-tuning corpus for models that turn noisy automatic-speech-recognition output into the written insertion a user intended. It contains 354,549 rows in official document-grouped 96/2/2 splits, 861.3 hours of processed 16 kHz speech, and 8.48M target-side loss tokens in 355 Parquet shards (49.25 GiB). This is not an ordinary transcription dataset. The model sees document context plus an ASR hypothesis and… See the full description on the dataset page: https://huggingface.co/datasets/johnbean393/fluid-2-sft-asr.audiotext-generation100K<n<1M0 likes10 downloads2mo agoHugging Face22Svngoku /kikongo-bible-asr-embeddings Kikongo Bible Embeddings This dataset is a version of the kikongo-bible-asr dataset. I used the cohere-emdbed-v3 model to produce the embeddings. texttext-generationn<1K1 likes9 downloads2y agoHugging Face23SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes9 downloads5mo agoHugging Face24vrclc /ASR-REF-PRED texttext-generation1K<n<10K1 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.