datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias.
Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay.
Description
All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.post-ocr-correction
Synthetic OCR Correction Dataset
This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction.
Description
To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.dfm12-dala-pl-correction
dfm12-dala-pl-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.sanskrit-ocr-post-correction\
A Benchmark and Dataset for Post-OCR text correction in Sanskrit.
This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script.
It includes:
- Train/Validation/Test splits with OCR text and corrected ground truth
- An out-of-domain test set of 500 sentences
- Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayadfm12-dala-is-correction
dfm12-dala-is-correction
Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status.
Only completed kept decisions with all three scores at least 4 are included, after deterministic gates.
Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations.
Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.text_error_correction文本纠错的相关数据
chinese_text_correction
Dataset Card
中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。
Repository: shibing624/pycorrector
Dataset Summary
拼写纠错数据
lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2
ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data
medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.g1_desktop_organize_table_correction
g1_desktop_organize_table_correction
Bimanual tabletop teleoperation on a Unitree G1 with Inspire dexterous hands,
in LeRobot v2.1 format.
Task — Follow the human's corrective gesture and hand the indicated object to the human.
Episodes
192
Frames
67405 (37.4 min @ 30 fps)
Episode length
268–519 frames (median 346)
State / action
26-D / 26-D
Cameras
2 × 640×480
State and action
Both vectors are 26-D and share the same layout:
0– 6… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/g1_desktop_organize_table_correction.Chinese-High-School-Chemistry-Correction-Dataset
Chinese-High-School-Chemistry-Correction-Dataset
一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集
1. 数据集缘起
为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。
2. 数据来源
普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式
该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。
3. 数据格式(JSONL)
每行一条记录,可直接用于 Hugging Face datasets 库:
{"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.grammar-correction
grammar-correction
Dataset Summary
The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset,
derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction.
It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models.
Dataset Structure
Train set: 100 000 entries
Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.rollout_stackblocks_iter2_correctionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/rollout_stackblocks_iter2_correction.policy-correction-c30-2026-09-12-13This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_piper_bimanual",
"total_episodes": 235,
"total_frames": 153515,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:235"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-c30-2026-09-12-13.stepwise_correction_sft_turn2_promptsim_double_insert_zheyuan_correction_0820policy-correction-2026-09-09This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_piper_bimanual",
"total_episodes": 186,
"total_frames": 305747,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:186"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-09.Punjabi-Gurmukhi-Grammar-Correction-Corpus
⚠️ Provenance note (2026-10-06). Sentence pairs were created with AI assistance following standard Punjabi grammar guidelines; they have not been reviewed by a professional linguist.
ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0)
👨💻 Research & Engineering Lead
Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio)
GitHub Profile:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.rollout_systematic_correction_smovla_3cam_blue_cube_orange_tray_20260825_130654This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_systematic_correction_smovla_3cam_blue_cube_orange_tray_20260825_130654.youtube_caption_corrections
Dataset Card for YouTube Caption Corrections
Dataset Summary
This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.300ep_blue_cube_orange_box_with_correctionsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/300ep_blue_cube_orange_box_with_corrections.HIL_corrections-only_2epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/grahamwichhh/HIL_corrections-only_2ep.rollout_correction_smolvla_3cam_blue_cube_orange_tray_20260824_205123This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_correction_smolvla_3cam_blue_cube_orange_tray_20260824_205123.fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.sim_double_insert_zheyuan_correction_0821rollout_picknplace_dagger_corrections_20260911_102054This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/xulla/rollout_picknplace_dagger_corrections_20260911_102054.rollout_stackblocks_iter1_correctionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/rollout_stackblocks_iter1_correction.policy-correction-2026-09-09-hgdagger-20fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_piper_bimanual",
"total_episodes": 158,
"total_frames": 15945,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:158"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-09-hgdagger-20fps.policy-correction-2026-09-14This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "agilex_piper_bimanual",
"total_episodes": 60,
"total_frames": 43460,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-14.rollout_1-100ep_smolvla_200ep_blue_cube_orange_box_corrections_20260819_191244This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_1-100ep_smolvla_200ep_blue_cube_orange_box_corrections_20260819_191244.
