Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay. Description All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.tabular10K<n<100K135 likes1.3k downloads1y agoHugging Face02jeanflop /post-ocr-correction Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.text1M<n<10M2 likes441 downloads2y agoHugging Face03schneiderkamplab /dfm12-dala-pl-correction dfm12-dala-pl-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.texttext-generation100K<n<1M0 likes427 downloads15d agoHugging Face04qualcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes338 downloads1d agoHugging Face05schneiderkamplab /dfm12-dala-is-correction dfm12-dala-is-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.texttext-generation100K<n<1M0 likes311 downloads15d agoHugging Face06shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes287 downloads2y agoHugging Face07agentlans /grammar-correction grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.texttext-classification100K<n<1M11 likes258 downloads2y agoHugging Face08ZihanC /stepwise_correction_sft_turn2_prompttext1K<n<10K0 likes204 downloads1y agoHugging Face09Nam-toon-studio /Punjabi-Gurmukhi-Grammar-Correction-Corpus ⚠️ Provenance note (2026-10-06). Sentence pairs were created with AI assistance following standard Punjabi grammar guidelines; they have not been reviewed by a professional linguist. ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus ☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Engineering Lead Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.texttext-generation1K<n<10K1 likes161 downloads1d agoHugging Face10community-datasets /youtube_caption_corrections Dataset Card for YouTube Caption Corrections Dataset Summary This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.textother10K<n<100K8 likes159 downloads2y agoHugging Face11greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes153 downloads2mo agoHugging Face12nrl-ai /vn-spell-correction-eval-real vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.texttext-generationn<1K0 likes152 downloads5mo agoHugging Face13hamsaai /AUTOSTT-ENG-correctionstextn<1K0 likes136 downloads19d agoHugging Face14AbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes127 downloads3mo agoHugging Face15Lots-of-LoRAs /task590_amazonfood_summary_correction_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task590_amazonfood_summary_correction_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task590_amazonfood_summary_correction_classification.texttext-generation1K<n<10K0 likes114 downloads2y agoHugging Face16Alexander-Usov /ru-asr-spell-correctiontext1K<n<10K0 likes99 downloads13d agoHugging Face17muzaffercky /kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models, like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections. The source videos are documented in the source.txt file. Usage from datasets import load_dataset dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train") print(dataset) textn<1K1 likes97 downloads1y agoHugging Face18neuripsedtracksub /ego-mistake-corrections Ego Mistake Corrections Benchmark (Ego-MC-Bench) Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/neuripsedtracksub/ego-mistake-corrections.textvideo-text-to-textn<1K0 likes96 downloads8d agoHugging Face19Metric-AI /fleurs-corrections FLEURS Armenian Reference Corrections This private dataset contains manual reference corrections used by ArmBench ASR for the Armenian (hy_am) test split of google/fleurs. Only rows whose reviewed transcription differs from the original FLEURS raw_transcription are included. Audio is not duplicated; use recording_id or file_name to join these rows to the upstream FLEURS test split. Fields recording_id: stable recording identifier, equal to the WAV filename stem.… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/fleurs-corrections.textn<1K0 likes95 downloads2mo agoHugging Face20Lots-of-LoRAs /task587_amazonfood_polarity_correction_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task587_amazonfood_polarity_correction_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task587_amazonfood_polarity_correction_classification.texttext-generation1K<n<10K0 likes90 downloads2y agoHugging Face21ajaysri /lego_stack_4x4_regular_plus_correction_frontdot_backwrist_heatmap_history_dynamic_lerobot_v3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "ur5_wsg50_lego_stack", "total_episodes": 202, "total_frames": 80912, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:202" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_4x4_regular_plus_correction_frontdot_backwrist_heatmap_history_dynamic_lerobot_v3.tabularrobotics10K<n<100K0 likes86 downloads4mo agoHugging Face22sbussiso /synthetic-self-correction-and-thinking-samples Self Correction and Thinking A seed library for training language models to reason with self-correction. Teaches three reasoning behaviors -- catching your own errors, verifying correct answers, and rejecting false doubts -- across four domains, three difficulty tiers, and three reasoning modes. Also includes multi-turn user-correction conversations where the user actively corrects or challenges the assistant. The structure at a glance graph TB… See the full description on the dataset page: https://huggingface.co/datasets/sbussiso/synthetic-self-correction-and-thinking-samples.imagetext-generation1K<n<10K0 likes84 downloads2mo agoHugging Face23coung21 /vi-spelling-correction Vietnamese Spelling Correction Dataset This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models. The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus. Dataset Structure The dataset is divided into training and testing sets: Train: 880,575 examples Test: 97,842 examples Data Fields source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.texttext-generation100K<n<1M1 likes81 downloads9mo agoHugging Face24google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes80 downloads3y agoHugging Face25Str1nger /ru-asr-spell-correctiontext1K<n<10K0 likes80 downloads11d agoHugging Face26Khamoon /asr-spell-correction-rutext1K<n<10K0 likes77 downloads28d agoHugging Face27latkes /self-consistency-correction-exp23-correlation-discrimination self-consistency-correction — Experiments 2 + 3 (combined) Paper: Correcting Generator Scores via Self-Consistency §7.3 (correlation) and §7.4 (discrimination). 20 hand-crafted cases with known equivalence classes covering geography, literature, science, history, and pop culture. Each case has a correct equivalence class containing 2-6 surface-form paraphrases plus 2 singleton wrong-answer classes (mix of rare and common string frequencies). Equivalence classes are hardcoded, so… See the full description on the dataset page: https://huggingface.co/datasets/latkes/self-consistency-correction-exp23-correlation-discrimination.tabularn<1K7 likes76 downloads6mo agoHugging Face28nguyenthanhasia /vsec-vietnamese-spell-correction VSEC: Vietnamese Spell Correction Dataset Dataset Description VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.texttext-generation1K<n<10K5 likes75 downloads1y agoHugging Face29fdemelo /spelling-correction-french-news Spelling correction dataset (French) This dataset is generated by transforming/corrupting sentences of a French news corpus provided by the University of Leipzig. The following transformations are applied to words in the sentences: concatenation of pairs of words swapping of neighboring letters in words insertion deletion replacement (by neighboring characters in AZERTY keyboard) Generation ./scripts/get_data.py -t news -y 2023 -s 10K ./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.text10K<n<100K1 likes74 downloads1y agoHugging Face30westenfelder /InterCode-Corrections Dataset Card for InterCode-Corrections This is a manually corrected version of the InterCode-Bash dataset, providing natural language prompts and Bash commands for the task of machine translation. Dataset Details Dataset Description This dataset contains corrections for errors in the InterCode-Bash dataset. corrections.csv contains annotations for each error. final.csv contains the updated dataset with the corrections applied. The corrected dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/westenfelder/InterCode-Corrections.texttranslationn<1K0 likes74 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.