Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes287 downloads2y agoHugging Face02AbhishekBhandari /Indic-post-ocr-correction Indic Contextual Post-OCR Correction Dataset Summary This dataset supports contextual post-OCR correction for Indic languages. Each example is a sentence-level triple consisting of: an OCR-generated sentence (noisy), the preceding sentence used as context, and the corrected sentence (ground truth). Hugging Face dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction Supported Tasks Post-OCR text correction… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Indic-post-ocr-correction.texttext-generation100K<n<1M1 likes127 downloads3mo agoHugging Face03muzaffercky /kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models, like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections. The source videos are documented in the source.txt file. Usage from datasets import load_dataset dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train") print(dataset) textn<1K1 likes97 downloads1y agoHugging Face04coung21 /vi-spelling-correction Vietnamese Spelling Correction Dataset This dataset contains 978,417 pairs of noisy (source) and clean (target) Vietnamese sentences, designed for training spelling correction models. The dataset was synthetically generated by injecting realistic noise into a clean Vietnamese corpus. Dataset Structure The dataset is divided into training and testing sets: Train: 880,575 examples Test: 97,842 examples Data Fields source: The text with injected errors (input).… See the full description on the dataset page: https://huggingface.co/datasets/coung21/vi-spelling-correction.texttext-generation100K<n<1M1 likes81 downloads9mo agoHugging Face05fdemelo /spelling-correction-french-news Spelling correction dataset (French) This dataset is generated by transforming/corrupting sentences of a French news corpus provided by the University of Leipzig. The following transformations are applied to words in the sentences: concatenation of pairs of words swapping of neighboring letters in words insertion deletion replacement (by neighboring characters in AZERTY keyboard) Generation ./scripts/get_data.py -t news -y 2023 -s 10K ./scripts/generate_dataset.py… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/spelling-correction-french-news.text10K<n<100K1 likes74 downloads1y agoHugging Face06yammdd /vietnamese-error-correction-corpus Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.text10K<n<100K0 likes74 downloads3mo agoHugging Face07chainiy /russian-spell-correctionstext1K<n<10K1 likes63 downloads20d agoHugging Face08s3h /arabic-grammar-correctionstext100K<n<1M7 likes47 downloads5y agoHugging Face09melsmm /spell-correction-ru Spell Correction RU — датасеты для коррекции ошибок в русском тексте Набор данных для обучения моделей исправления орфографических, пунктуационных и регистровых ошибок в русскоязычных текстах. Каждый пример — пара «правильный текст» → «текст с ошибкой». Датасет использовался для обучения модели melsmm/Spell-Corrector-RU-4B. 📦 Код генерации, ноутбуки и полное описание проекта: github.com/melsmm/llm-spell-corrector Состав Датасет содержит две конфигурации… See the full description on the dataset page: https://huggingface.co/datasets/melsmm/spell-correction-ru.texttext-generation1M<n<10M1 likes47 downloads4mo agoHugging Face10ClarusC64 /clinical-iatrogenic-self-correction-breakdown-analysis-v0.2 Clinical Iatrogenic Self Correction Breakdown Analysis v0.2 What this is A small dataset that tests one question: Can you detect when a clinical system is moving toward self-correction breakdown, not just carrying iatrogenic pressure? This repo focuses on iatrogenic self-correction failure. It models a system where: iatrogenic pressure may rise self-correction capacity may shrink feedback integrity may weaken compensatory fatigue may accumulate before overt collapse… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-iatrogenic-self-correction-breakdown-analysis-v0.2.tabulartext-classificationn<1K0 likes42 downloads7mo agoHugging Face11xlp100 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/xlp100/chinese_text_correction.text100K<n<1M0 likes41 downloads7mo agoHugging Face12seanghay /khmer-spelling-corrections Khmer Spelling Corrections Naturally occurring Khmer misspellings paired with the word the writer meant. The labels are not annotated, they are observed. Search sessions record the whole typing trajectory toward a single word, so when a user types something, fails, adjusts and lands on a real dictionary headword, the failed attempt and the headword form a correction pair produced by a real person under no instruction to make mistakes. Only pairs within two edits of the target… See the full description on the dataset page: https://huggingface.co/datasets/seanghay/khmer-spelling-corrections.tabulartext-generationn<1K2 likes40 downloads2mo agoHugging Face13pbevan11 /synthetic-ocr-correction-gpt4o Synthetic OCR Correction GPT-4o 10,000 pieces of news text from fancyzhx/ag_news with synthetically generated OCR mistakes. The purpose of this is to mimic corrupt text that has been transcribed with OCR from old newspapers, where there are often lot's of errors. See biglam/bnl_newspapers1841-1879 for example. By synthetically creating it, we have the true ground truth, meaning we can use this as a source of truth for finetuning. The corrupted text was generated using OpenAI's… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/synthetic-ocr-correction-gpt4o.text10K<n<100K6 likes39 downloads2y agoHugging Face14ambrosfitz /2k_grammar_correctionstexttext-generation1K<n<10K1 likes33 downloads1y agoHugging Face15holygleb /asr-correction-rutext1K<n<10K0 likes32 downloads14d agoHugging Face16u042e /asr-spell-correction-synthetictext1K<n<10K0 likes29 downloads14d agoHugging Face17Owishiboo /grammar-correctionBasically used in Correctness Chorus to train T5 model to predict grammar correction. text1K<n<10K2 likes28 downloads4y agoHugging Face18ClarusC64 /clinical-iatrogenic-self-correction-breakdown-analysis-v0.1What this dataset tests Whether an intelligence system can explainwhy a clinical system failed to self-correctduring an iatrogenic harm cascade. Required outputs failed interruption points recovery opportunities lost resilience gap index normalization of deviance flags authority gradient effects preventability score minimal system fix set Use case Third layer of the Iatrogenic Harm Cascade Library. tabulartabular-classificationn<1K0 likes27 downloads8mo agoHugging Face19ClarusC64 /quantum-error-correction-failure-v0.1 quantum-error-correction-failure-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum error correction regimes. Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction. The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency. Core stability idea Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.tabulartabular-classificationn<1K0 likes25 downloads5mo agoHugging Face20bmd1905 /error-correction-vitext100K<n<1M8 likes24 downloads4y agoHugging Face21linagora /synthetic_corrections Read the paper here Synthetic Corrections Data This corpus consists of 300 dialogues between Architect and Builder, inspired by the interactions found in the Minecraft Structured Dialogue Corpus. A single dialogue sample contains three Architect instructions, each of which requires Builder to construct a simple shape on the grid. Builder correctly constructs two out of the three shapes; the final Architect turn contains a Correction of the incorrect construction. The task… See the full description on the dataset page: https://huggingface.co/datasets/linagora/synthetic_corrections.textn<1K0 likes24 downloads1y agoHugging Face22huytd189 /japanese-grammar-correction Japanese Grammar Correction Dataset Description This dataset is designed to train language models to identify and correct a wide range of grammatical errors and stylistic issues in Japanese text. The data consists of pairs of incorrect and correct sentences, along with metadata that classifies the type of error and provides additional context. The dataset was created by both manual curation from discussions in Japanese learning communities and synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/huytd189/japanese-grammar-correction.texttext-generation1K<n<10K0 likes24 downloads1y agoHugging Face23Haradrick228 /gendl-hw1-asr-correction-rutext1K<n<10K0 likes23 downloads6d agoHugging Face24aranemini /central-kurdish-correction-table Central Kurdish Orthographic Correction Table This repository contains a correction table designed to standardize Central Kurdish (Sorani) script for speech and language technology applications. Orthographic variation is a major challenge in Kurdish NLP. Different spellings and writing conventions for the same words can introduce inconsistencies in ASR training, evaluation, and downstream NLP systems. This table provides mappings from non-standard or inconsistent forms to their… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-correction-table.texttext-generation10K<n<100K0 likes22 downloads2mo agoHugging Face25v0es /ru-proper-noun-spell-correctionstext1K<n<10K0 likes22 downloads1d agoHugging Face26bmd1905 /vi-error-correction-2.0text1M<n<10M1 likes21 downloads2y agoHugging Face27llllliuuy /Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction Title An Annotated Dataset for Hearing-Impaired Speech-to-Text Correction Abstract This dataset is specifically designed for the task of correcting speech to text errors in hearing-impaired individuals. It includes one hour of real speech files of hearing-impaired individuals, automatic speech recognition (ASR) output text, and manually corrected standard text. We searched for an hour of audio from hearing-impaired individuals to ensure that the voice was authentic and… See the full description on the dataset page: https://huggingface.co/datasets/llllliuuy/Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction.text1K<n<10K0 likes21 downloads6mo agoHugging Face28ClarusC64 /rscl-correction-quality-under-constraint-v0.1 What this dataset tests Whether a correction is both: correct compliant with constraints A model can fix the contentand still violate the rules. This dataset separates those cases. Why this exists Self-correction often fails by: fixing the answer but breaking the format complying with the format but keeping the error introducing new errors during repair This benchmark scores correction quality under constraint. Data format Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/rscl-correction-quality-under-constraint-v0.1.texttext-classificationn<1K0 likes20 downloads8mo agoHugging Face29bmd1905 /vi-error-correction-v2text100K<n<1M3 likes17 downloads2y agoHugging Face30floriandebaene /EmDComF_OCR_post-correctionOCR-to-gold parallel dataset for OCR post-correcting early modern Dutch, originating from EmDComF. text10K<n<100K2 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.