datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-self-correction-and-thinking-samples
Self Correction and Thinking
A seed library for training language models to reason with self-correction.
Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant.
The structure at a glance
graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.lego_stack_4x4_n100_corr100_postrelease4cm_correctionstart_rerender_lerobot_v7corrections_pick_place_black_king_jan_20Emakhuwa-Portuguese-OCR-post-correctionBibTeX:
The dataset paper was published in EMNLP 2024.
Please cite as:
@inproceedings{ali-etal-2024-building,
title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks",
author = "Ali, Felermino D. M. A. and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.Ocr_Post_Correctionlego_stack_4x4_n100_corr100_postrelease4cm_correctionstartframe_lerobot_v7lego_stack_4x4_n100_corr100_postrelease4cm_correctionparent_colorisolated_lerobot_v6receipt-ocr-correctionsomnidoc-ocr-correction-bench
OmniDoc OCR Correction Bench
A benchmark dataset for evaluating VLMs on OCR error correction and document-to-markdown formatting.
Overview
Each sample pairs a document image from OmniDocBench v1.5 with a prompt containing PaddleOCR-extracted markdown text. The task is to correct OCR errors and restore proper formatting using the source image as reference.
Dataset Structure
Field
Type
Description
prompt
string
System prompt with OCR-extracted markdown… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/omnidoc-ocr-correction-bench.jfk-ocr-correctionjfk-2025-tiny-ocr-no-correctionsearch_with_correction-trainmy-dental-privacy-correctionsoja-correction
