correction
Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias.
Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay.
Description
All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.post-ocr-correction
Synthetic OCR Correction Dataset
This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction.
Description
To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.dfm12-dala-pl-correction
dfm12-dala-pl-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.sanskrit-ocr-post-correction\
A Benchmark and Dataset for Post-OCR text correction in Sanskrit.
This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script.
It includes:
- Train/Validation/Test splits with OCR text and corrected ground truth
- An out-of-domain test set of 500 sentences
- Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayadfm12-dala-is-correction
dfm12-dala-is-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.
