Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes518 downloads1mo agoHugging Face02yuanzhuyun /asr-reference-set-eval-temp Temporary ASR evaluation audio Temporary public audio files used for hosted ASR evaluation. audio1K<n<10K0 likes358 downloads2mo agoHugging Face03GenSEC-LLM /SLT-Task2-Post-ASR-Speaker-Tagging Dataset Name: Dataset for ASR Speaker-Tagging Corrections (Speaker Diarization) Description This dataset is pairs of erroneous ASR output and speaker tagging, which are generated from a ASR system and speaker diarization system. Each source erroneous transcription is paired with human-annotated transcription, which has correct transcription and speaker tagging. SEGment-wise Long-form Speech Transcription annotation (SegLST), the file format used in the CHiME challenges… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task2-Post-ASR-Speaker-Tagging.tabular10K<n<100K2 likes281 downloads2y agoHugging Face04Edge0 /ark-asr-open-asr-leaderboard-results ARK-ASR Open ASR Leaderboard Results This dataset contains JSONL prediction manifests for AutoArk-AI/ARK-ASR-0.6B on hf-audio/open-asr-leaderboard public English short-form splits. These files are intended for Open ASR Leaderboard maintainer verification. Scoring summary from normalizer.eval_utils.score_results: Split WER RTFx ami/test 10.02 352.12 earnings22/test 9.77 331.88 gigaspeech/test 8.00 217.72 librispeech/test.clean 1.53 412.12 librispeech/test.other… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-open-asr-leaderboard-results.tabular10K<n<100K13 likes216 downloads4mo agoHugging Face05NightPrince /quran-asr-husary Quran ASR — Husary Muallim Dataset Description This dataset contains Quran recitation audio files by Sheikh Mahmoud Khalil Al-Husary at 16 kHz sampling rate, with Arabic transcriptions including diacritics. Dataset Structure Audio files: Stored in audio/ folder (e.g., audio/001_001.wav) Data file: manifest.json (NeMo format) Columns: audio_filepath: Path to audio file text: Arabic transcription with diacritics duration: Audio duration in seconds speaker:… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/quran-asr-husary.audioautomatic-speech-recognition1K<n<10K0 likes214 downloads7mo agoHugging Face06elihoole /asrs-aviation-reports Dataset Card for ASRS Aviation Incident Reports Dataset Summary This dataset collects 47,723 aviation incident reports published in the Aviation Safety Reporting System (ASRS) database maintained by NASA. Supported Tasks and Leaderboards 'summarization': Dataset can be used to train a model for abstractive and extractive summarization. The model performance is measured by how high the output summary's ROUGE score for a given narrative account of an aviation… See the full description on the dataset page: https://huggingface.co/datasets/elihoole/asrs-aviation-reports.textsummarization10K<n<100K11 likes211 downloads4y agoHugging Face07Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes202 downloads4mo agoHugging Face08whyismydininghallonfire /orig-plus-asr-tamil-clean orig-plus-asr-tamil-clean Combined ASR dataset built from: albagon/til26-asr-split (orig rows) whyismydininghallonfire/asr-tamil-clean (asr_tamil_clean rows) Audio paths are namespaced under each split to avoid filename collisions: audio/orig/... audio/asr_tamil_clean/... Each row keeps key, audio, transcript, and language, with an added source_dataset field. Counts: train: 3595 orig + 891 asr_tamil_clean = 4486 validation: 899 orig + 224 asr_tamil_clean = 1123 audio1K<n<10K0 likes191 downloads4mo agoHugging Face09benderrodriguez /hebrew-asr-vn Hebrew ASR three-source training dataset Derived from ivrit.ai crowd-transcribe-v5, crowd-recital and knesset-committees. Original data and transcripts are credited to ivrit.ai and its contributors. Pinned revisions and preparation rules are in metadata/sources.json and metadata/preparation-config.json. VoxKnesset is excluded by user decision. Source/split Clips Hours crowd-recital/test 1,557 1.071 crowd-recital/train 45,372 33.258 crowd-recital/validation 1,033… See the full description on the dataset page: https://huggingface.co/datasets/benderrodriguez/hebrew-asr-vn.tabularautomatic-speech-recognition1M<n<10M0 likes188 downloads27d agoHugging Face10thethanksforthegod /quran-asr-mega-corpustabular10K<n<100K1 likes173 downloads1mo agoHugging Face11KUOCHENG /ASR_NeurIPSn<1K0 likes144 downloads1y agoHugging Face12uzinfocom-edu-ai /uzbek-asr-curated-701h Uzbek ASR Curated Dataset (701 hours) A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation. Dataset Description Language Uzbek (Latin script with okina ʻ) Total utterances 337,920 Total duration ~701 hours Audio format 16 kHz mono WAV (PCM_16) Manifest format NeMo JSONL Splits train (94%) / val (3%) / test (3%) Splits Split Utterances Hours Train 317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.audioautomatic-speech-recognition100K<n<1M1 likes137 downloads4mo agoHugging Face13tterumiimurett1 /agentic-asrgated Agentic ASR Public consolidated audio and ASR result dataset for the OSWorld and WildClawBench benchmark families. Layout osworld/: synthetic raw/colloquial speech, human recordings, DNS-noise pairs, task images, ASR results, and reports. wildclawbench/: 60 formal colloquialized prompts, synthetic speech, 20 synthetic ASR condition tables, and ten-participant human recordings. task0_template derivatives are excluded. metadata/conditions.jsonl: model, variant… See the full description on the dataset page: https://huggingface.co/datasets/tterumiimurett1/agentic-asr.audio10K<n<100K0 likes131 downloads19d agoHugging Face14JesseParvess /book_snippets_asrtextn<1K0 likes129 downloads5y agoHugging Face15danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes120 downloads11mo agoHugging Face16SaarAI /asr-benchmark-outputsgated SaarAI ASR Benchmark Outputs Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets. files: 570 utterances: 4615160 languages: 7 models: 50 Layout data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl index.jsonl # one record per file (language, split, model, rows, sha256, ...) index.csv Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.tabularautomatic-speech-recognition1M<n<10M1 likes113 downloads5d agoHugging Face17bosonai /asr-eval-results-r3atabular10K<n<100K0 likes87 downloads4mo agoHugging Face18google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes80 downloads3y agoHugging Face19ImFastAsBlitz /gendl-hw1-asr-spell-ollama Исправление ошибок распознавания речи — GenDL HW1 Сгенерированный мой датаест Поля Поле Содержимое input Исходный текст, в котором могут быть ошибки распознавания. output Целевой исправленный текст. Файлы asr_spell_train.jsonl — основной датасет: одна JSON-запись на строку. asr_spell_train.csv — копия тех же данных в CSV. Загрузка from datasets import load_dataset dataset = load_dataset(… See the full description on the dataset page: https://huggingface.co/datasets/ImFastAsBlitz/gendl-hw1-asr-spell-ollama.texttext-generation1K<n<10K0 likes76 downloads14d agoHugging Face20Meddies /Meddies-ASR-Benchgated Meddies ASR Bench This public dataset repository contains the validated 10 full audio chunks used for the MOSS + Gemini 3.8 Flash smoke benchmark. Contents data/full_10_chunks/audio/: ten full MP3 input chunks. data/full_10_chunks/moss/: the ten source MOSS-Diarize-Transcribe JSON records. data/full_10_chunks/manifest.json: chunk paths, offsets, durations, and embedded MOSS text. results/: the downloaded Gemini 3.8 hosted run artifacts, including the ten native… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/Meddies-ASR-Bench.audioautomatic-speech-recognitionn<1K1 likes74 downloads12d agoHugging Face21QuechuaBase /asr-ser-quechua-collao-embeddings ASR-SER embeddings for Quechua Collao This repository contains embeddings only. It does not contain raw audio. These embeddings were extracted for ASR-to-SER transfer experiments on the Quechua Collao emotional speech corpus. The release is intended for reproducible feature-extraction experiments and downstream analysis. Dataset contents One PyTorch tensor per utterance stored as an embedding file under embeddings/ A sanitized metadata table describing utterance… See the full description on the dataset page: https://huggingface.co/datasets/QuechuaBase/asr-ser-quechua-collao-embeddings.tabularaudio-classificationn<1K0 likes68 downloads3mo agoHugging Face22sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes67 downloads24d agoHugging Face23ketav /parakeet-hindi-asr Parakeet Hindi-English Bilingual ASR Fine-tuning NVIDIA Parakeet TDT 0.6B for bilingual Hindi-English automatic speech recognition. Quick Start # Download pip install huggingface_hub huggingface-cli download ketav/parakeet-hindi-asr --repo-type dataset --local-dir ./parakeet-hindi-asr # Install dependencies pip install nemo_toolkit[asr] bitsandbytes sentencepiece # Train (after updating paths in config) cd parakeet-hindi-asr/scripts python ft_0.6B_hi_v3.py… See the full description on the dataset page: https://huggingface.co/datasets/ketav/parakeet-hindi-asr.textautomatic-speech-recognition100K<n<1M0 likes66 downloads9mo agoHugging Face24Kartal-Ol /azb-asr-corpustext100K<n<1M1 likes65 downloads4mo agoHugging Face25annmakarova /russian-asr-correctionstext1K<n<10K0 likes65 downloads13d agoHugging Face26lawmaluki /KambaBench-ASR KambaBench-ASR Status: v0.0 — scaffold. No evaluation audio or gold transcriptions have been finalized yet. An open, leakage-controlled, reproducible evaluation benchmark for Kamba (Kikamba, kam) automatic speech recognition (ASR). KambaBench-ASR is designed to provide a common evaluation standard for Kamba speech-recognition systems. The benchmark is intended to be model-agnostic: any Kamba ASR system, whether based on Whisper, MMS, Omnilingual ASR, Parakeet, or another… See the full description on the dataset page: https://huggingface.co/datasets/lawmaluki/KambaBench-ASR.textautomatic-speech-recognitionn<1K1 likes62 downloads2mo agoHugging Face27woongstar /ko-finance-asr-corrections ko-finance-asr-corrections Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube. 210 pairs mined from 2,391 videos of auto-captions across 47 channels totalling 1,080.1 hours Each pair carries how often the term was mangled and how often it was said correctly, plus verification provenance. 한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답 표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다. What makes it different No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.tabulartext-generationn<1K0 likes62 downloads29d agoHugging Face28rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes57 downloads1mo agoHugging Face29zadterishi /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/zadterishi/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes56 downloads15d agoHugging Face30abnajlae /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes54 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.