Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes84 downloads2mo agoHugging Face02ajaysri /lego_stack_4x4_n100_corr100_postrelease4cm_correctionstart_rerender_lerobot_v7imagen<1K0 likes46 downloads3mo agoHugging Face03thewisp /corrections_pick_place_black_king_jan_20imagen<1K0 likes35 downloads9mo agoHugging Face04LIACC /Emakhuwa-Portuguese-OCR-post-correctionBibTeX: The dataset paper was published in EMNLP 2024. Please cite as: @inproceedings{ali-etal-2024-building, title = "Building Resources for Emakhuwa: Machine Translation and News Classification Benchmarks", author = "Ali, Felermino D. M. A. and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen, Yun-Nung", booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-OCR-post-correction.imagetranslationn<1K0 likes20 downloads2y agoHugging Face05Alijeff1214 /Ocr_Post_Correctionimagen<1K0 likes19 downloads2y agoHugging Face06ajaysri /lego_stack_4x4_n100_corr100_postrelease4cm_correctionstartframe_lerobot_v7imagen<1K0 likes19 downloads3mo agoHugging Face07ajaysri /lego_stack_4x4_n100_corr100_postrelease4cm_correctionparent_colorisolated_lerobot_v6imagen<1K0 likes18 downloads3mo agoHugging Face08patelparthm2000 /receipt-ocr-correctionsimage10K<n<100K0 likes17 downloads3mo agoHugging Face09andynoodles /omnidoc-ocr-correction-bench OmniDoc OCR Correction Bench A benchmark dataset for evaluating VLMs on OCR error correction and document-to-markdown formatting. Overview Each sample pairs a document image from OmniDocBench v1.5 with a prompt containing PaddleOCR-extracted markdown text. The task is to correct OCR errors and restore proper formatting using the source image as reference. Dataset Structure Field Type Description prompt string System prompt with OCR-extracted markdown… See the full description on the dataset page: https://huggingface.co/datasets/andynoodles/omnidoc-ocr-correction-bench.imageimage-to-text1K<n<10K1 likes12 downloads5mo agoHugging Face10zdeng314 /jfk-ocr-correctionimagen<1K0 likes7 downloads1y agoHugging Face11zdeng314 /jfk-2025-tiny-ocr-no-correctionimagen<1K0 likes5 downloads1y agoHugging Face12ajaysri /search_with_correction-trainimagen<1K0 likes5 downloads1y agoHugging Face13omaster /my-dental-privacy-correctionimagen<1K0 likes4 downloads1y agoHugging Face14Guguinhaxd /soja-correctionimagen<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.