yammdd/vietnamese-error-correction-corpus
Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.
Data Summary
The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.
• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.
• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.
• Data Characteristics: The dataset includes common Vietnamese text errors such as missing diacritics, spelling mistakes, teencode, abbreviations, and informal language.
• Data Quality: As the annotations are generated automatically and sourced from social media, the dataset may contain noise or imperfect corrections.
• Ethical Note: The dataset is not intended to include offensive content or to target any individual or organization.
Data Format
The dataset is organized into two columns:
• Input: Noisy Vietnamese text, including missing diacritics, spelling errors, teencode, and informal variants.
• Target: Corrected Vietnamese text with proper diacritics, spelling, grammar, and normalized informal expressions.
