Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cyberfish /text_error_correction文本纠错的相关数据 1 likes254 downloads5y agoHugging Face02google /red_ace_asr_error_detection_and_correction RED-ACE Dataset Summary This dataset can be used to train and evaluate ASR Error Detection or Correction models. It was introduced in the RED-ACE paper (Gekhman et al, 2022). The dataset contains ASR outputs on the LibriSpeech corpus (Panayotov et al., 2015) with annotated transcription errors. Dataset Details The LibriSpeech corpus was decoded using Google Cloud Speech-to-Text API, with the default and video models. The word-level confidence was enabled… See the full description on the dataset page: https://huggingface.co/datasets/google/red_ace_asr_error_detection_and_correction.textautomatic-speech-recognition100K<n<1M6 likes98 downloads3y agoHugging Face03muzaffercky /kurdish-kurmanji-grammar-error-correctionThis dataset is for developing and evaluating grammatical error correction (GEC) models, like Grammarly, for Kurdish Kurmanji. Incorrect sentences were manually collected from YouTube comment sections of Kurdish videos and X(Twitter) and Muzaffer Cıkay added their corrections. The source videos are documented in the source.txt file. Usage from datasets import load_dataset dataset = load_dataset("muzaffercky/kurdish-kurmanji-typo-correction", split="train") print(dataset) textn<1K1 likes89 downloads1y agoHugging Face048uBob /text_error_correction文本纠错的相关数据 0 likes81 downloads2mo agoHugging Face05yammdd /vietnamese-error-correction-corpus Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.text10K<n<100K0 likes71 downloads3mo agoHugging Face06sajjadiba /urdu-asr-error-correction-data Urdu ASR Generative Error Correction Dataset This dataset contains paired training and testing data for post-ASR error correction in Urdu. Dataset Details Language: Urdu (ur) Task: ASR Error Correction License: CC BY-NC 4.0 Dataset Structure The dataset consists of parallel text pairs containing raw ASR transcripts generated by Whisper-large-v3-turbo alongside their corresponding target corrections (pseudo-gold). train.jsonl / train.csv:… See the full description on the dataset page: https://huggingface.co/datasets/sajjadiba/urdu-asr-error-correction-data.text1K<n<10K0 likes66 downloads22d agoHugging Face07sarayusapa /Grammar_Error_Correctiontext100K<n<1M0 likes58 downloads1y agoHugging Face08SyntheticLogic-Labs /python-runtime-verified-error-correction Python Runtime-Verified Error Correction Dataset 🐍⚡ Overview Production-grade synthetic dataset of Python code errors with runtime-verified corrections. Each sample contains broken code, the actual runtime error, and a guaranteed-working fix validated through execution. Unlike traditional synthetic datasets, every correction is verified by actually running the code in an isolated environment—eliminating hallucinations and ensuring real-world applicability.… See the full description on the dataset page: https://huggingface.co/datasets/SyntheticLogic-Labs/python-runtime-verified-error-correction.texttext-generation1K<n<10K0 likes41 downloads9mo agoHugging Face09puppalupa /ru-asr-error-correctiontext1K<n<10K0 likes41 downloads3d agoHugging Face10slone /bak_ocr_error_correction_2022 Dataset Card for "bak_ocr_error_correction_2022" More Information needed text10K<n<100K0 likes30 downloads3y agoHugging Face11p208p2002 /zhtw-sentence-error-correction 中文錯字糾正資料集 由規則與字典自維基百科產生的錯誤糾正資料集。 包含錯誤類型:隨機錯字、近似音錯字、缺字錯誤、冗字錯誤。 資料集使用函式庫: p208p2002/zh-mistake-text-gen 子集 alpha: 95%錯誤,5%不變。單句中可能有多個錯誤。 beta: 50%錯誤,50%不變。單句中僅有一個錯誤。 gamma: 100%錯誤。單句中可能有多個錯誤。 text100K<n<1M5 likes30 downloads3y agoHugging Face12schneiderkamplab /dfm11-folketingets-dokumenter-error-correction DFM11 Folketingets Dokumenter Error Correction This dataset is the fully audited DFM11 replacement for schneiderkamplab/dfm10-folketingets-dokumenter-error-correction. Every retained input was generated from its target using 1-8 declared synthetic OCR substitutions. Deterministic text-quality filtering was followed by a task-aware Gemma 4 audit of all 2,548,956 surviving rows; 63,109 audit rejections were removed and 2,485,847 rows remain. Rows contain messages in… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-folketingets-dokumenter-error-correction.text1M<n<10M0 likes27 downloads1mo agoHugging Face13bmd1905 /error-correction-vitext100K<n<1M8 likes26 downloads4y agoHugging Face14ClarusC64 /quantum-error-correction-failure-v0.1 quantum-error-correction-failure-v0.1 What this dataset does This dataset evaluates whether models can detect instability in quantum error correction regimes. Each row represents a simplified quantum computing scenario where logical qubits are protected using error correction. The task is to determine whether the correction mechanism remains stable or fails due to noise and correction latency. Core stability idea Quantum error correction works by detecting and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/quantum-error-correction-failure-v0.1.tabulartabular-classificationn<1K0 likes26 downloads5mo agoHugging Face15Karthi02 /grammatical_error_correction0 likes24 downloads1y agoHugging Face16bmd1905 /vi-error-correction-v2text100K<n<1M3 likes19 downloads2y agoHugging Face17Karthi02 /grammatical-error-correction0 likes19 downloads1y agoHugging Face18bmd1905 /vi-error-correction-2.0text1M<n<10M1 likes18 downloads2y agoHugging Face19shahidul034 /error_correction_model_dataset_raw Dataset Card for "error_correction_model_dataset" More Information needed text1M<n<10M0 likes16 downloads4y agoHugging Face20Shoriful025 /quantum_error_correction_telemetrytabularn<1K0 likes16 downloads9mo agoHugging Face21SPEAK-PP /synthetic-error-generated-spelling-correction-dataset-100ktext10K<n<100K0 likes16 downloads6mo agoHugging Face22sumitaryal /nepali_grammatical_error_correctiontext1M<n<10M0 likes15 downloads2y agoHugging Face23vosap52 /robot-error-correction-tr-v1 Robot Error Correction TR v1 This dataset focuses on failure detection and corrective behavior in embodied AI systems. Unlike standard instruction datasets, each sample represents: an incorrect real-world outcome a corrective decision The goal is improving humanoid robot autonomy and reliability in real environments. Capabilities trained: self-correction safety awareness environment feedback handling recovery planning textroboticsn<1K0 likes13 downloads8mo agoHugging Face24schneiderkamplab /dfm10-folketingets-dokumenter-error-correction dfm10-folketingets-dokumenter-error-correction Audited folketingets-dokumenter-error-correction tasks derived from Folketing documents. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 13 Rows: 3,105,440 Category: Danish transformation Upstream material Rigsarkivet handover 14004 / Folketinget Processing The complete generated task… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-folketingets-dokumenter-error-correction.0 likes12 downloads1mo agoHugging Face25reverendish /advanced-math-error-correctiontext100K<n<1M1 likes9 downloads5mo agoHugging Face26kilicai /turkish-sft-error-correction-10k kilicai/turkish-sft-error-correction-10k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error-correction-10k') text10K<n<100K1 likes8 downloads5mo agoHugging Face27kilicai /turkish-sft-error_correction_20k kilicai/turkish-sft-error_correction_20k Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('kilicai/turkish-sft-error_correction_20k') text10K<n<100K0 likes8 downloads5mo agoHugging Face28duyle2408 /vi-error-correctiontext1K<n<10K0 likes4 downloads1y agoHugging Face29duyle2408 /vi-error-correction-2text10K<n<100K0 likes4 downloads1y agoHugging Face30duyle2408 /vi-error-correction-4text1K<n<10K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.