datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.Indic-post-ocr-correction
Indic Contextual Post-OCR Correction
Dataset Summary
This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of:
an OCR-generated sentence (noisy),
the preceding sentence used as context, and
the corrected sentence (ground truth).
Hugging Face dataset page:
https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction
Supported Tasks
Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models,
like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected
from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections.
The source videos are documented in the source.txt file.
Usage
from datasets import load_dataset
dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train")
print(dataset)
vi-spelling-correction
Vietnamese Spelling Correction Dataset
This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models.
The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus.
Dataset Structure
The dataset is divided into training and testing sets:
Train: 880,575 examples
Test: 97,842 examples
Data Fields
source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.spelling-correction-french-news
Spelling correction dataset (French)
This dataset is generated by transforming/corrupting sentences of a French news corpus
provided by the University of Leipzig.
The following transformations are applied to words in the sentences:
concatenation of pairs of words
swapping of neighboring letters in words
insertion
deletion
replacement (by neighboring characters in AZERTY keyboard)
Generation
./scripts/get_data.py -t news -y 2023 -s 10K
./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.vietnamese-error-correction-corpus
Data Summary
The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.
• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.
• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.
• Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.russian-spell-correctionsarabic-grammar-correctionsspell-correction-ru
Spell Correction RU — датасеты для коррекции ошибок в русском тексте
Набор данных для обучения моделей исправления орфографических, пунктуационных и регистровых ошибок в русскоязычных текстах.
Каждый пример — пара «правильный текст» → «текст с ошибкой». Датасет использовался для обучения модели melsmm/Spell-Corrector-RU-4B.
📦 Код генерации, ноутбуки и полное описание проекта: github.com/melsmm/llm-spell-corrector
Состав
Датасет содержит две конфигурации… See the full description on the dataset page: https://huggingface.co/datasets/melsmm/spell-correction-ru.clinical-iatrogenic-self-correction-breakdown-analysis-v0.2
Clinical Iatrogenic Self Correction Breakdown Analysis v0.2
What this is
A small dataset that tests one question:
Can you detect when a clinical system is moving toward self-correction breakdown, not just carrying iatrogenic pressure?
This repo focuses on iatrogenic self-correction failure.
It models a system where:
iatrogenic pressure may rise
self-correction capacity may shrink
feedback integrity may weaken
compensatory fatigue may accumulate before overt collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-iatrogenic-self-correction-breakdown-analysis-v0.2.chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/xlp100/chinese_text_correction.khmer-spelling-corrections
Khmer Spelling Corrections
Naturally occurring Khmer misspellings paired with the word the writer meant.
The labels are not annotated, they are observed. Search sessions record the
whole typing trajectory toward a single word, so when a user types something,
fails, adjusts and lands on a real dictionary headword, the failed attempt and
the headword form a correction pair produced by a real person under no
instruction to make mistakes.
Only pairs within two edits of the target… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-spelling-corrections.synthetic-ocr-correction-gpt4o
Synthetic OCR Correction GPT-4o
10,000 pieces of news text from fancyzhx/ag_news with synthetically generated OCR mistakes.
The purpose of this is to mimic corrupt text that has been transcribed with OCR from old newspapers, where there are often lot's of errors. See biglam/bnl_newspapers1841-1879 for example. By synthetically creating it, we have the true ground truth, meaning we can use this as a source of truth for finetuning.
The corrupted text was generated using OpenAI's… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/synthetic-ocr-correction-gpt4o.2k_grammar_correctionsasr-correction-ruasr-spell-correction-syntheticgrammar-correctionBasically used in Correctness Chorus to train T5 model to predict grammar correction.
clinical-iatrogenic-self-correction-breakdown-analysis-v0.1What this dataset tests
Whether an intelligence system can explainwhy a clinical system failed to self-correctduring an iatrogenic harm cascade.
Required outputs
failed interruption points
recovery opportunities lost
resilience gap index
normalization of deviance flags
authority gradient effects
preventability score
minimal system fix set
Use case
Third layer of the Iatrogenic Harm Cascade Library.
quantum-error-correction-failure-v0.1
quantum-error-correction-failure-v0.1
What this dataset does
This dataset evaluates whether models can detect instability in quantum error correction regimes.
Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction.
The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency.
Core stability idea
Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.error-correction-visynthetic_corrections
Read the paper here
Synthetic Corrections Data
This corpus consists of 300 dialogues between Architect and Builder, inspired by the interactions found in the Minecraft Structured Dialogue Corpus. A single dialogue sample contains three Architect instructions, each of which requires Builder to construct a simple shape on the grid. Builder correctly constructs two out of the three shapes; the final Architect turn contains a Correction of the incorrect construction. The task… See the full description on the dataset page: https://huggingface.co/datasets/linagora/synthetic_corrections.japanese-grammar-correction
Japanese Grammar Correction Dataset
Description
This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text.
The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context.
The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.gendl-hw1-asr-correction-rucentral-kurdish-correction-table
Central Kurdish Orthographic Correction Table
This repository contains a correction table designed to standardize Central Kurdish (Sorani) script for speech and language technology applications.
Orthographic variation is a major challenge in Kurdish NLP. Different spellings and writing conventions for the same words can introduce inconsistencies in ASR training, evaluation, and downstream NLP systems.
This table provides mappings from non-standard or inconsistent forms to their… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-correction-table.ru-proper-noun-spell-correctionsvi-error-correction-2.0Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction
Title
An Annotated Dataset for Hearing-Impaired Speech-to-Text Correction
Abstract
This dataset is specifically designed for the task of correcting speech to text errors in hearing-impaired individuals.
It includes one hour of real speech files of hearing-impaired individuals, automatic speech recognition (ASR) output text, and manually corrected standard text.
We searched for an hour of audio from hearing-impaired individuals to ensure that the voice was authentic and… See the full description on the dataset page: https://huggingface.co/datasets/llllliuuy/Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction.rscl-correction-quality-under-constraint-v0.1
What this dataset tests
Whether a correction is both:
correct
compliant with constraints
A model can fix the contentand still violate the rules.
This dataset separates those cases.
Why this exists
Self-correction often fails by:
fixing the answer but breaking the format
complying with the format but keeping the error
introducing new errors during repair
This benchmark scores correction quality under constraint.
Data format
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/rscl-correction-quality-under-constraint-v0.1.vi-error-correction-v2EmDComF_OCR_post-correctionOCR-to-gold parallel dataset for OCR post-correcting early modern Dutch, originating from EmDComF.
