Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01schneiderkamplab /dfm12-dala-pl-correction dfm12-dala-pl-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.texttext-generation100K<n<1M0 likes427 downloads15d agoHugging Face02qualcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes338 downloads2d agoHugging Face03schneiderkamplab /dfm12-dala-is-correction dfm12-dala-is-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.texttext-generation100K<n<1M0 likes311 downloads15d agoHugging Face04agentlans /grammar-correction grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.texttext-classification100K<n<1M11 likes258 downloads2y agoHugging Face05Nam-toon-studio /Punjabi-Gurmukhi-Grammar-Correction-Corpus ⚠️ Provenance note (2026-10-06). Sentence pairs were created with AI assistance following standard Punjabi grammar guidelines; they have not been reviewed by a professional linguist. ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus ☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Engineering Lead Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.texttext-generation1K<n<10K1 likes161 downloads1d agoHugging Face06greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes153 downloads2mo agoHugging Face07nrl-ai /vn-spell-correction-eval-real vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.texttext-generationn<1K0 likes152 downloads5mo agoHugging Face08hamsaai /AUTOSTT-ENG-correctionstextn<1K0 likes136 downloads19d agoHugging Face09neuripsedtracksub /ego-mistake-corrections Ego Mistake Corrections Benchmark (Ego-MC-Bench) Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-mistake-corrections.textvideo-text-to-textn<1K0 likes96 downloads8d agoHugging Face10sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes84 downloads2mo agoHugging Face11google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes80 downloads3y agoHugging Face12westenfelder /InterCode-Corrections Dataset Card for InterCode-Corrections This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation. Dataset Details Dataset Description This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.texttranslationn<1K0 likes74 downloads2y agoHugging Face13nrl-ai /vn-spell-correction-train nrl-ai/vn-spell-correction-train 459,478 (noisy, clean) Vietnamese training pairs for fine-tuning a seq2seq spell-correction model. Each row: {"input": "<noisy>", "target": "<clean>"} Both fields are NFC-normalized. How it was built Clean side: same 500K register-balanced mix as nrl-ai/vn-diacritic-train — 350K Vietnamese Wikipedia (CC-BY-SA-4.0, hirine/wikipedia-vietnamese-1M296K-dataset) + 150K NFC-fixed Vietnamese news (CC-BY-4.0, tmnam20/Vietnamese-News-dedup).… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-train.texttext-generation100K<n<1M0 likes68 downloads5mo agoHugging Face14sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes67 downloads24d agoHugging Face15schneiderkamplab /dfm12-dala-fo-correction dfm12-dala-fo-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-fo-correction.texttext-generation100K<n<1M0 likes66 downloads16d agoHugging Face16schneiderkamplab /dfm12-dala-nn-correction dfm12-dala-nn-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-nn-correction.texttext-generation100K<n<1M0 likes65 downloads15d agoHugging Face17annmakarova /russian-asr-correctionstext1K<n<10K0 likes65 downloads13d agoHugging Face18woongstar /ko-finance-asr-corrections ko-finance-asr-corrections Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube. 210 pairs mined from 2,391 videos of auto-captions across 47 channels totalling 1,080.1 hours Each pair carries how often the term was mangled and how often it was said correctly, plus verification provenance. 한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답 표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다. What makes it different No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.tabulartext-generationn<1K0 likes62 downloads1mo agoHugging Face19schneiderkamplab /dfm12-dala-nb-correction dfm12-dala-nb-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-nb-correction.texttext-generation100K<n<1M0 likes61 downloads15d agoHugging Face20nrl-ai /vn-spell-correction-eval nrl-ai/vn-spell-correction-eval Vietnamese spell-correction evaluation grid: 4 source registers × 2 noise levels = 8 splits, 2,098 (noisy, clean) sentence pairs total. Each pair is {"input": "<noisy>", "target": "<clean>"}. Both sides are NFC-normalized. The clean target is the same sentence used as the target in nrl-ai/vn-diacritic-eval — spell correction is a strict superset of diacritic restoration, so we reuse the same registers-balanced corpus. Splits Two noise… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval.texttext-generation1K<n<10K0 likes60 downloads5mo agoHugging Face21torinriley /spell-correction Spell-Check Dataset This dataset consists of pairs of misspelled words and their corresponding correctly spelled words, designed for training and evaluating character-level spelling correction models. It is particularly useful for tasks such as: Spelling correction Character-level sequence-to-sequence modeling Error detection and correction in text Each data point in the dataset contains: misspelled: A misspelled version of a word. correct: The corrected spelling of the word.… See the full description on the dataset page: https://huggingface.co/datasets/torinriley/spell-correction.text10K<n<100K2 likes52 downloads2y agoHugging Face22dougalldeepmind /2026-09-15-dataset-refresh-correction-audit Dataset refresh correction audit; not a training release field value experiment Zero-new-API correction of the incomplete refresh: 40 net independent exclusion reversals and one lossless completed-review parsing recovery. Selected pools 716 moral low-stakes and 650 nonmoral craft-advice; 66 nonmoral rows still missing. Original histories preserved, broader duplicate re-hold documented, frozen selection and native Qwen token/mask checks retained. Four saved-answer… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-correction-audit.tabularn<1K0 likes52 downloads25d agoHugging Face23schneiderkamplab /dfm12-dala-sv-correction dfm12-dala-sv-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-sv-correction.texttext-generation100K<n<1M0 likes50 downloads16d agoHugging Face24True2456 /gemma4-onpolicy-student-corrections Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill). Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.texttext-generation1K<n<10K0 likes48 downloads3mo agoHugging Face25marcelone /text-correction_collection Human Samples These samples contains contains human-written sentences produced during language learning practice, combined with AI-based grammatical verification and correction. The original sentences were written by language learners who often did not know whether their sentences were correct or incorrect. These authentic learner inputs capture a wide range of natural mistakes, such as spelling, syntax, word choice, and structure errors. Synthetic Samples These… See the full description on the dataset page: https://huggingface.co/datasets/marcelone/text-correction_collection.texttext-generation1K<n<10K0 likes42 downloads11mo agoHugging Face26SyntheticLogic-Labs /python-runtime-verified-error-correction Python Runtime-Verified Error Correction Dataset 🐍⚡ Overview Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution. Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.texttext-generation1K<n<10K0 likes42 downloads9mo agoHugging Face27True2456 /gemma4-onpolicy-50topics-2000-corrections Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.texttext-generation1K<n<10K0 likes38 downloads3mo agoHugging Face28arthurdubrou /Bird_explained_corrections Dataset Card for Dataset Name This dataset is truncated textn<1K0 likes34 downloads3y agoHugging Face29KokunoYumeto /stacks-project-corrections Stacks project corrections 2,301 textual corrections to the Stacks project, the open reference on algebraic geometry, each pinned to the official revision a04446e57ec1. The errors were found by GPT-5.6 (OpenAI) while translating the Stacks project, then collected, checked and worked through by GPT-6 Astra (OpenAI), and integrated into the AI-integrated Stacks project, a fork where the review evidence for each correction is kept. These are AI-reviewed corrections. They are not… See the full description on the dataset page: https://huggingface.co/datasets/KokunoYumeto/stacks-project-corrections.text1K<n<10K0 likes34 downloads3d agoHugging Face30protonx-models /text-correction-validationtext100K<n<1M11 likes31 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.