Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muzaffercky /kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models, like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections. The source videos are documented in the source.txt file. Usage from datasets import load_dataset dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train") print(dataset) textn<1K1 likes97 downloads1y agoHugging Face02google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes80 downloads3y agoHugging Face03yammdd /vietnamese-error-correction-corpus Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.text10K<n<100K0 likes74 downloads3mo agoHugging Face04sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes67 downloads25d agoHugging Face05puppalupa /ru-asr-error-correctiontext1K<n<10K0 likes43 downloads6d agoHugging Face06SyntheticLogic-Labs /python-runtime-verified-error-correction Python Runtime-Verified Error Correction Dataset 🐍⚡ Overview Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution. Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.texttext-generation1K<n<10K0 likes42 downloads9mo agoHugging Face07sarayusapa /Grammar_Error_Correctiontext100K<n<1M0 likes40 downloads1y agoHugging Face08slone /bak_ocr_error_correction_2022 Dataset Card for "bak_ocr_error_correction_2022" More Information needed text10K<n<100K0 likes28 downloads3y agoHugging Face09p208p2002 /zhtw-sentence-error-correction 中文錯字糾正資料集 由規則與字典自維基百科產生的錯誤糾正資料集。 包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。 資料集使用函式庫: p208p2002/zh-mistake-text-gen 子集 alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。 beta: 50%錯誤,50%不變。單句中僅有一個錯誤。 gamma: 100%錯誤。單句中可能有多個錯誤。 text100K<n<1M5 likes26 downloads3y agoHugging Face10ClarusC64 /quantum-error-correction-failure-v0.1 quantum-error-correction-failure-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum error correction regimes. Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction. The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency. Core stability idea Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.tabulartabular-classificationn<1K0 likes25 downloads5mo agoHugging Face11bmd1905 /error-correction-vitext100K<n<1M8 likes24 downloads4y agoHugging Face12schneiderkamplab /dfm11-folketingets-dokumenter-error-correction DFM11 Folketingets Dokumenter Error Correction This dataset is the fully audited DFM11 replacement for schneiderkamplab/dfm10-folketingets-dokumenter-error-correction. Every retained input was generated from its target using 1-8 declared synthetic OCR substitutions. Deterministic text-quality filtering was followed by a task-aware Gemma 4 audit of all 2,548,956 surviving rows; 63,109 audit rejections were removed and 2,485,847 rows remain. Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.text1M<n<10M0 likes23 downloads1mo agoHugging Face13bmd1905 /vi-error-correction-2.0text1M<n<10M1 likes21 downloads2y agoHugging Face14bmd1905 /vi-error-correction-v2text100K<n<1M3 likes17 downloads2y agoHugging Face15Shoriful025 /quantum_error_correction_telemetrytabularn<1K0 likes16 downloads9mo agoHugging Face16shahidul034 /error_correction_model_dataset_raw Dataset Card for "error_correction_model_dataset" More Information needed text1M<n<10M0 likes15 downloads4y agoHugging Face17SPEAK-PP /synthetic-error-generated-spelling-correction-dataset-100ktext10K<n<100K0 likes15 downloads6mo agoHugging Face18vosap52 /robot-error-correction-tr-v1 Robot Error Correction TR v1 This dataset focuses on failure detection and corrective behavior in embodied AI systems. Unlike standard instruction datasets, each sample represents: an incorrect real-world outcome a corrective decision The goal is improving humanoid robot autonomy and reliability in real environments. Capabilities trained: self-correction safety awareness environment feedback handling recovery planning textroboticsn<1K0 likes14 downloads8mo agoHugging Face19sumitaryal /nepali_grammatical_error_correctiontext1M<n<10M0 likes13 downloads2y agoHugging Face20reverendish /advanced-math-error-correctiontext100K<n<1M1 likes10 downloads5mo agoHugging Face21kilicai /turkish-sft-error-correction-10k kilicai/turkish-sft-error-correction-10k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error-correction-10k') text10K<n<100K1 likes8 downloads5mo agoHugging Face22kilicai /turkish-sft-error_correction_20k kilicai/turkish-sft-error_correction_20k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error_correction_20k') text10K<n<100K0 likes8 downloads5mo agoHugging Face23duyle2408 /vi-error-correctiontext1K<n<10K0 likes5 downloads1y agoHugging Face24duyle2408 /vi-error-correction-2text10K<n<100K0 likes4 downloads1y agoHugging Face25duyle2408 /vi-error-correction-4text1K<n<10K0 likes4 downloads1y agoHugging Face26duyle2408 /vi-error-correction-5text10K<n<100K0 likes4 downloads1y agoHugging Face27duyle2408 /vi-error-correction-super-small-upper-2text1K<n<10K0 likes4 downloads1y agoHugging Face28duyle2408 /vi-error-correction-6text10K<n<100K0 likes3 downloads1y agoHugging Face29duyle2408 /vi-error-correction-9text10K<n<100K0 likes3 downloads1y agoHugging Face30duyle2408 /vi-error-correction-small-10text1K<n<10K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.