Team Ai
Datasetpublic

yammdd/vietnamese-error-correction-corpus

Data Summary The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets. • Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text. • Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs. • Data… See the full description on the dataset page: https://huggingface.co/datasets/yammdd/vietnamese-error-correction-corpus.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes74downloads
Dataset Card

Data Summary

The model is trained on a Vietnamese text error correction dataset constructed from real-world noisy inputs. The dataset contains approximately 70,000 sentence pairs and is split into training, validation, and test sets.

• Data Source: Crawled Vietnamese social media comments, reflecting informal and user-generated text.

• Annotation Method: Automatically labeled using a large language model, which generates corrected versions of noisy inputs.

• Data Characteristics: The dataset includes common Vietnamese text errors such as missing diacritics, spelling mistakes, teencode, abbreviations, and informal language.

• Data Quality: As the annotations are generated automatically and sourced from social media, the dataset may contain noise or imperfect corrections.

• Ethical Note: The dataset is not intended to include offensive content or to target any individual or organization.

Data Format

The dataset is organized into two columns:

• Input: Noisy Vietnamese text, including missing diacritics, spelling errors, teencode, and informal variants.

• Target: Corrected Vietnamese text with proper diacritics, spelling, grammar, and normalized informal expressions.