Team Ai
Datasetpublic

tejasta/Kannada-English-Parallel-Corpus

Dataset Card for Kannada-English Parallel Corpus Dataset Description This dataset provides a high-quality, curated collection of parallel English-Kannada sentence pairs. It is designed to address the critical data scarcity in low-resource language modeling. This corpus is intended to facilitate advancements in Neural Machine Translation (NMT), cross-lingual transfer learning, and instruction-tuning for Large Language Models (LLMs) to better serve Kannada speakers.… See the full description on the dataset page: https://huggingface.co/datasets/tejasta/Kannada-English-Parallel-Corpus.

sourceHugging Facemitupdated 3mo agoView on Hugging Face
7likes10downloads
Dataset Card

Dataset Card for Kannada-English Parallel Corpus

Dataset Description

This dataset provides a high-quality, curated collection of parallel English-Kannada sentence pairs. It is designed to address the critical data scarcity in low-resource language modeling. This corpus is intended to facilitate advancements in Neural Machine Translation (NMT), cross-lingual transfer learning, and instruction-tuning for Large Language Models (LLMs) to better serve Kannada speakers.

  • —Curated by: Tejas Tumakuru Ashok
  • —Language(s): English (en), Kannada (kn)
  • —License: MIT

Uses

Direct Use

This dataset is intended for:

  • —Neural Machine Translation (NMT): Training and evaluating models to translate between English and Kannada.
  • —LLM Fine-tuning: Instruction-tuning models to improve their ability to understand and generate fluent, contextually accurate Kannada.
  • —Cross-Lingual Transfer Learning: Developing models that leverage English-language pre-training to improve performance on Kannada-specific tasks.

Out-of-Scope Use

  • —This dataset should not be used to generate medical, legal, or financial advice without rigorous expert human-in-the-loop verification.
  • —It is not intended for use in systems that generate or propagate hate speech, discrimination, or harmful content.

Dataset Structure

The dataset consists of parallel sentence pairs formatted for ease of use in Hugging Face datasets pipelines.

  • —`english`: The source text in English.
  • —`kannada`: The corresponding translation in Kannada.

Dataset Creation

Curation Rationale

Kannada is a resource-poor language in the context of modern AI. Despite its large number of speakers, the lack of high-quality, digitized parallel data hampers the development of localized AI tools. This dataset was created to bridge this digital divide, ensuring that Kannada speakers can benefit from advancements in Natural Language Processing.

Source Data

The data was collected from open-source web crawls, public government corpora.

Data Collection and Processing
  • —Cleaning: The data underwent rigorous filtering to remove duplicates, non-Kannada/English characters, and improperly aligned translations.
  • —Normalization: Text was normalized to ensure standard Unicode representation for Kannada characters.
  • —Validation: Validation was performed using algorithmic evaluation, agentic loop validation, and manual spot checks. Kannada, my native language, was used for the manual validation.

Personal and Sensitive Information

This dataset does not contain personally identifiable information (PII). All data has been anonymized or sourced from public-domain repositories.

Bias, Risks, and Limitations

  • —Linguistic Bias: As with most parallel corpora, translations may reflect the dialectal or formal nature of the source text.
  • —Coverage: The dataset may lack coverage for highly technical or colloquial Kannada domains.
  • —Recommendation: Users should evaluate the performance of models trained on this dataset against a gold-standard evaluation set to ensure it meets the specific requirements of their application.