Team Ai
20 results

correction

PleIAs /Post-OCR-CorrectionPost-OCR correction is a large corpus of 1 billion words containing original texts with a varying number of OCR mistakes and an experimental multilingual post-OCR correction output created by Pleias. Generation of Post-OCR correction was performed using HPC resources from GENCI–IDRIS (Grant 2023-AD011014736) on Jean-Zay. Description All the texts come from collections integrated into Common Corpus, the largest open corpus for pretraining previously released by Pleias on HuggingFace.… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Post-OCR-Correction.tabular10K<n<100K135 likes1.3k downloads1y agoHugging Facejeanflop /post-ocr-correction Synthetic OCR Correction Dataset This dataset is a synthetic dataset generated for post-OCR correction tasks. It contains over 2,000,000 rows of French text pairs and follows the Croissant format. It is designed to train small language models (LLMs) for text correction. Description To ensure the dataset closely resembles OCR-malformed texts, we applied various transformations randomly. This approach helps avoid the LLM identifying specific patterns and encourages it to… See the full description on the dataset page: https://huggingface.co/datasets/jeanflop/post-ocr-correction.text1M<n<10M2 likes441 downloads2y agoHugging Faceschneiderkamplab /dfm12-dala-pl-correction dfm12-dala-pl-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-pl-correction.texttext-generation100K<n<1M0 likes427 downloads15d agoHugging Facequalcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes338 downloads1d agoHugging Faceacomquest /sanskrit-ocr-post-correction\ A Benchmark and Dataset for Post-OCR text correction in Sanskrit. This dataset contains manually post-edited OCR data for Sanskrit texts in Devanagari script. It includes: - Train/Validation/Test splits with OCR text and corrected ground truth - An out-of-domain test set of 500 sentences - Source texts from classical Sanskrit works including Brahmasutra Bhashyam, Grahalaghava, and Goladhyayatext-classification100K<n<1M2 likes334 downloads1y agoHugging Faceschneiderkamplab /dfm12-dala-is-correction dfm12-dala-is-correction Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant target indices are preserved.… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-dala-is-correction.texttext-generation100K<n<1M0 likes311 downloads15d agoHugging Face