Team Ai
Datasetpublic

CoRover/hplt

HPLT Indic Language Corpus This repository contains selected Indic-language data from the HPLT Monolingual Dataset 3.0, prepared and hosted by CoRover for large-scale generative language model pretraining and multilingual NLP research. The dataset contains raw HPLT data for multiple Indian languages and scripts. Languages The repository currently contains the following language/script datasets: Language Language Code Script Directory Bengali ben Bengali… See the full description on the dataset page: https://huggingface.co/datasets/CoRover/hplt.

sourceHugging Facecc0-1.0updated 1mo agoView on Hugging Face
0likes370downloads
Dataset Card

HPLT Indic Language Corpus

This repository contains selected Indic-language data from the HPLT Monolingual Dataset 3.0, prepared and hosted by CoRover for large-scale generative language model pretraining and multilingual NLP research.

The dataset contains raw HPLT data for multiple Indian languages and scripts.

Languages

The repository currently contains the following language/script datasets:

LanguageLanguage CodeScriptDirectory
BengalibenBengaliben_Beng_raw
GujaratigujGujaratiguj_Gujr_raw
HindihinDevanagarihin_Deva_raw
KannadakanKannadakan_Knda_raw
MalayalammalMalayalammal_Mlym_raw
MarathimarDevanagarimar_Deva_raw
ManipurimniBengalimni_Beng_raw
OdiaoryOdiaory_Orya_raw
PunjabipanGurmukhipan_Guru_raw
TamiltamTamiltam_Taml_raw
TelugutelTelugutel_Telu_raw
UrduurdArabicurd_Arab_raw

Intended Use

This dataset is intended for:

  • —Generative language model pretraining
  • —Causal language modeling
  • —Multilingual language model training
  • —Indic language model training
  • —Continued pretraining
  • —Text generation
  • —NLP research
  • —Tokenizer development
  • —Cross-lingual language modeling
  • —Low-resource language research

The primary objective of this repository is to provide large-scale Indic-language data suitable for training and developing generative AI models.

Indic Language Coverage

The repository focuses on major Indian languages represented in their commonly used scripts:

  • —Bengali
  • —Gujarati
  • —Hindi
  • —Kannada
  • —Malayalam
  • —Marathi
  • —Manipuri
  • —Odia
  • —Punjabi
  • —Tamil
  • —Telugu
  • —Urdu

These datasets can be combined to train multilingual models with support for multiple Indic languages.

HPLT

HPLT (High Performance Language Technologies) is a large-scale multilingual web corpus project.

HPLT Monolingual Dataset 3.0 provides data for a large number of language-script combinations and is designed for training large language models and other NLP systems.

The source data is primarily derived from large-scale web crawls, including Common Crawl and Internet Archive data.

The HPLT processing pipeline includes operations such as:

  • —Text extraction
  • —Language identification
  • —Quality filtering
  • —Deduplication
  • —Document processing
  • —Language/script classification

Data Quality

The HPLT dataset contains quality-related information and is distributed in shards that can be selected according to the requirements of a particular training pipeline.

For production model training, users may perform additional:

  • —Language filtering
  • —Quality filtering
  • —Document deduplication
  • —Domain filtering
  • —Text normalization
  • —Script validation
  • —Toxicity and safety filtering

Recommended Usage

For language model pretraining, it is recommended to:

  1. 1.Decompress the required HPLT shards.
  2. 2.Extract the document text.
  3. 3.Apply language and script validation.
  4. 4.Remove duplicate documents.
  5. 5.Apply quality filtering.
  6. 6.Normalize the text where appropriate.
  7. 7.Tokenize using the target model tokenizer.
  8. 8.Convert the data into the required training format.

For multilingual training, language sampling or weighting can be applied to balance languages with different corpus sizes.

Source Dataset

This repository is derived from:

HPLT Monolingual Dataset 3.0

The original HPLT project provides information about the dataset, processing pipeline, language coverage, statistics, and terms of use.

License

The HPLT project releases the dataset packaging under CC0.

The underlying text originates from web sources and may be subject to rights, licenses, terms of service, or other restrictions applicable to the original content.

Users are responsible for determining whether their intended use of the data complies with applicable laws and the rights associated with the underlying source material.

Citation

If you use the HPLT dataset in research or model development, please cite the HPLT project.

bibtex
@misc{hplt,
  title = {HPLT Monolingual Datasets},
  author = {HPLT},
  url = {https://hplt-project.org/datasets/v3.0}
}

Acknowledgements

We acknowledge the HPLT (High Performance Language Technologies) project for creating and releasing the multilingual corpus used as the source for this repository.

This repository is maintained by CoRover and provides selected Indic-language HPLT data for generative AI, language model pretraining, and NLP research.

Disclaimer

This repository contains web-derived data. CoRover does not claim ownership of the underlying third-party content contained in the source corpus.

Users should independently evaluate the data for copyright, licensing, privacy, safety, quality, and other applicable requirements before using it for model training or downstream applications.