datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias.
Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay.
Description
All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.post-ocr-correction
Synthetic OCR Correction Dataset
This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction.
Description
To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.dfm12-dala-pl-correction
dfm12-dala-pl-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.dfm12-dala-is-correction
dfm12-dala-is-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.grammar-correction
grammar-correction
Dataset Summary
The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset,
derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction.
It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models.
Dataset Structure
Train set: 100 000 entries
Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.stepwise_correction_sft_turn2_promptPunjabi-Gurmukhi-Grammar-Correction-Corpus
⚠️ Provenance note (2026-10-06). Sentence pairs were created with AI assistance following standard Punjabi grammar guidelines; they have not been reviewed by a professional linguist.
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.youtube_caption_corrections
Dataset Card for YouTube Caption Corrections
Dataset Summary
This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.AUTOSTT-ENG-correctionsIndic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.task590_amazonfood_summary_correction_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task590_amazonfood_summary_correction_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task590_amazonfood_summary_correction_classification.ru-asr-spell-correctionkurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models,
like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected
from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections.
The source videos are documented in the source.txt file.
Usage
from datasets import load_dataset
dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train")
print(dataset)
ego-mistake-corrections
Ego Mistake Corrections Benchmark (Ego-MC-Bench)
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-mistake-corrections.fleurs-corrections
FLEURS Armenian Reference Corrections
This private dataset contains manual reference corrections used by
ArmBench ASR for the
Armenian (hy_am) test split of google/fleurs.
Only rows whose reviewed transcription differs from the original FLEURS
raw_transcription are included. Audio is not duplicated; use recording_id
or file_name to join these rows to the upstream FLEURS test split.
Fields
recording_id: stable recording identifier, equal to the WAV filename stem.… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/fleurs-corrections.task587_amazonfood_polarity_correction_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task587_amazonfood_polarity_correction_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task587_amazonfood_polarity_correction_classification.lego_stack_4x4_regular_plus_correction_frontdot_backwrist_heatmap_history_dynamic_lerobot_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 202,
"total_frames": 80912,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:202"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_4x4_regular_plus_correction_frontdot_backwrist_heatmap_history_dynamic_lerobot_v3.synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.vi-spelling-correction
Vietnamese Spelling Correction Dataset
This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models.
The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus.
Dataset Structure
The dataset is divided into training and testing sets:
Train: 880,575 examples
Test: 97,842 examples
Data Fields
source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.red_ace_asr_error_detection_and_correction
RED-ACE
Dataset Summary
This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022).
The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors.
Dataset Details
The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models.
The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.ru-asr-spell-correctionasr-spell-correction-ruself-consistency-correction-exp23-correlation-discrimination
self-consistency-correction — Experiments 2 + 3 (combined)
Paper: Correcting Generator Scores via Self-Consistency §7.3 (correlation) and §7.4 (discrimination).
20 hand-crafted cases with known equivalence classes covering geography, literature, science, history, and pop culture. Each case has a correct equivalence class containing 2-6 surface-form paraphrases plus 2 singleton wrong-answer classes (mix of rare and common string frequencies). Equivalence classes are hardcoded, so… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp23-correlation-discrimination.vsec-vietnamese-spell-correction
VSEC: Vietnamese Spell Correction Dataset
Dataset Description
VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.spelling-correction-french-news
Spelling correction dataset (French)
This dataset is generated by transforming/corrupting sentences of a French news corpus
provided by the University of Leipzig.
The following transformations are applied to words in the sentences:
concatenation of pairs of words
swapping of neighboring letters in words
insertion
deletion
replacement (by neighboring characters in AZERTY keyboard)
Generation
./scripts/get_data.py -t news -y 2023 -s 10K
./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.InterCode-Corrections
Dataset Card for InterCode-Corrections
This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation.
Dataset Details
Dataset Description
This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.
