Team Ai
Datasetpublic

icfoss/English_Malayalam_Translation_Human_annotated

English-Malayalam Government Parallel Corpus Synth This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text. Machine Translation Notice All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/English_Malayalam_Translation_Human_annotated.

sourceHugging Faceotherupdated 26d agoView on Hugging Face
0likes51downloads
Dataset Card

English-Malayalam Government Parallel Corpus Synth

This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text.

Machine Translation Notice

All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional review.

Dataset Details

  • —Languages: English (en) and Malayalam (ml)
  • —Rows: 58,174 aligned pairs
  • —Format: CSV
  • —Source file: translation.zip / translation.csv
  • —Text domain: Government and administrative prose
  • —Data type: Synthetic / machine-translated parallel text

Schema

ColumnTypeDescription
englishstringEnglish source sentence
malayalamstringMalayalam aligned sentence

Intended Uses

  • —English-Malayalam machine translation
  • —Domain adaptation for Malayalam language models
  • —Parallel corpus filtering and alignment research
  • —Government-domain terminology extraction

Data Processing

The original CSV was normalized by:

  • —retaining only the English and Malayalam text columns
  • —dropping empty unnamed columns
  • —retaining rows where both sides are non-empty
  • —trimming leading and trailing whitespace

Limitations

The corpus is synthetic machine-translated text and may contain noisy alignments, administrative boilerplate, OCR artifacts, translation artifacts, and domain-specific terminology. Licensing was not inferred automatically; users should verify reuse permissions.