Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PleIAs /Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay. Description All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.tabular10K<n<100K135 likes1.3k downloads1y agoHugging Face02jeanflop /post-ocr-correction Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.text1M<n<10M2 likes441 downloads2y agoHugging Face03schneiderkamplab /dfm12-dala-pl-correction dfm12-dala-pl-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.texttext-generation100K<n<1M0 likes427 downloads15d agoHugging Face04qualcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes338 downloads1d agoHugging Face05acomquest /sanskrit-ocr-post-correction\ A Benchmark and Dataset for Post-OCR text correction in Sanskrit. This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script. It includes: - Train/Validation/Test splits with OCR text and corrected ground truth - An out-of-domain test set of 500 sentences - Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayatext-classification100K<n<1M2 likes334 downloads1y agoHugging Face06schneiderkamplab /dfm12-dala-is-correction dfm12-dala-is-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.texttext-generation100K<n<1M0 likes311 downloads15d agoHugging Face07Cyberfish /text_error_correction文本纠错的相关数据 1 likes289 downloads5y agoHugging Face08shibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes287 downloads2y agoHugging Face09DaoyuanZhu /g1_desktop_organize_table_correction g1_desktop_organize_table_correction Bimanual tabletop teleoperation on a Unitree G1 with Inspire dexterous hands, in LeRobot v2.1 format. Task — Follow the human's corrective gesture and hand the indicated object to the human. Episodes 192 Frames 67405 (37.4 min @ 30 fps) Episode length 268–519 frames (median 346) State / action 26-D / 26-D Cameras 2 × 640×480 State and action Both vectors are 26-D and share the same layout: 0– 6… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/g1_desktop_organize_table_correction.tabularrobotics10K<n<100K0 likes262 downloads1mo agoHugging Face10liushuaiqian /Chinese-High-School-Chemistry-Correction-Dataset Chinese-High-School-Chemistry-Correction-Dataset 一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集 1. 数据集缘起 为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。 2. 数据来源 普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式 该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。 3. 数据格式(JSONL) 每行一条记录,可直接用于 Hugging Face datasets 库: {"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.question-answering1K<n<10K4 likes259 downloads1y agoHugging Face11agentlans /grammar-correction grammar-correction Dataset Summary The grammar-correction dataset is a refined subset of the liweili/c4_200m dataset, derived from Google's C4_200M Synthetic Dataset for Grammatical Error Correction. It contains sentence pairs where the input is ungrammatical and the output is grammatical, making it suitable for training grammatical error correction (GEC) models. Dataset Structure Train set: 100 000 entries Validation set: 25 000 entries… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/grammar-correction.texttext-classification100K<n<1M11 likes258 downloads2y agoHugging Face12lilkm /rollout_stackblocks_iter2_correctionThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/rollout_stackblocks_iter2_correction.tabularrobotics10K<n<100K0 likes230 downloads3mo agoHugging Face13yashgoyal0110 /policy-correction-c30-2026-09-12-13This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "agilex_piper_bimanual", "total_episodes": 235, "total_frames": 153515, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:235" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-c30-2026-09-12-13.tabularrobotics100K<n<1M0 likes219 downloads26d agoHugging Face14ZihanC /stepwise_correction_sft_turn2_prompttext1K<n<10K0 likes204 downloads1y agoHugging Face15huzheyuan /sim_double_insert_zheyuan_correction_0820videon<1K0 likes181 downloads3mo agoHugging Face16yashgoyal0110 /policy-correction-2026-09-09This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "agilex_piper_bimanual", "total_episodes": 186, "total_frames": 305747, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:186" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-09.tabularrobotics100K<n<1M1 likes177 downloads1mo agoHugging Face17Nam-toon-studio /Punjabi-Gurmukhi-Grammar-Correction-Corpus ⚠️ Provenance note (2026-10-06). Sentence pairs were created with AI assistance following standard Punjabi grammar guidelines; they have not been reviewed by a professional linguist. ੴ Punjabi (Gurmukhi) Grammatical Error Correction Corpus ☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਆਕਰਣ ਸ਼ੁੱਧੀ ਅਤੇ ਸੁਧਾਰ ਡਾਟਾਸੈੱਟ (v1.0) 👨‍💻 Research & Engineering Lead Creator & Architect: Gurpreet Singh Dhillon (Nam-toon Studio) GitHub Profile:… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-Gurmukhi-Grammar-Correction-Corpus.texttext-generation1K<n<10K1 likes161 downloads1d agoHugging Face18makermods /rollout_systematic_correction_smovla_3cam_blue_cube_orange_tray_20260825_130654This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_systematic_correction_smovla_3cam_blue_cube_orange_tray_20260825_130654.tabularrobotics1K<n<10K0 likes160 downloads2mo agoHugging Face19community-datasets /youtube_caption_corrections Dataset Card for YouTube Caption Corrections Dataset Summary This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.textother10K<n<100K8 likes159 downloads2y agoHugging Face20makermods /300ep_blue_cube_orange_box_with_correctionsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/300ep_blue_cube_orange_box_with_corrections.tabularrobotics10K<n<100K0 likes158 downloads2mo agoHugging Face21grahamwichhh /HIL_corrections-only_2epThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/grahamwichhh/HIL_corrections-only_2ep.tabularroboticsn<1K0 likes156 downloads1mo agoHugging Face22makermods /rollout_correction_smolvla_3cam_blue_cube_orange_tray_20260824_205123This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_correction_smolvla_3cam_blue_cube_orange_tray_20260824_205123.tabularrobotics1K<n<10K0 likes154 downloads2mo agoHugging Face23greghavens /fable-5-coding-and-debugging-traces-synthetic-corrections Model Synthetic Corrections 1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset. Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.tabulartext-generationn<1K0 likes153 downloads2mo agoHugging Face24nrl-ai /vn-spell-correction-eval-real vn-spell-correction-eval-real Out-of-distribution evaluation corpus for Vietnamese spell-correction models — 150 hand-curated (noisy, clean) pairs sampled from real VN error sources, not generated by nom.text.noise. This is the test set we use to verify a spell-correction model generalises beyond its own synthetic training distribution. A model that scores 95 % on nom-vn's synthetic eval grid and 60 % on this set is overfit to the noise generator. Splits Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.texttext-generationn<1K0 likes152 downloads5mo agoHugging Face25huzheyuan /sim_double_insert_zheyuan_correction_0821videon<1K0 likes148 downloads3mo agoHugging Face26xulla /rollout_picknplace_dagger_corrections_20260911_102054This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/xulla/rollout_picknplace_dagger_corrections_20260911_102054.tabularrobotics10K<n<100K0 likes146 downloads29d agoHugging Face27lilkm /rollout_stackblocks_iter1_correctionThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/rollout_stackblocks_iter1_correction.tabularrobotics10K<n<100K0 likes141 downloads3mo agoHugging Face28yashgoyal0110 /policy-correction-2026-09-09-hgdagger-20fpsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "agilex_piper_bimanual", "total_episodes": 158, "total_frames": 15945, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:158" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-09-hgdagger-20fps.tabularrobotics10K<n<100K0 likes139 downloads29d agoHugging Face29yashgoyal0110 /policy-correction-2026-09-14This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "agilex_piper_bimanual", "total_episodes": 60, "total_frames": 43460, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:60" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/yashgoyal0110/policy-correction-2026-09-14.tabularrobotics10K<n<100K0 likes138 downloads25d agoHugging Face30makermods /rollout_1-100ep_smolvla_200ep_blue_cube_orange_box_corrections_20260819_191244This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/makermods/rollout_1-100ep_smolvla_200ep_blue_cube_orange_box_corrections_20260819_191244.tabularrobotics1K<n<10K0 likes137 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.