Team Ai
Datasetpublic

hyper-efficient-system-llc/bosnian-corpus-v1

Bosnian Corpus v1.0 (cleaned) This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks. The canonical release of the corpus is archived on Zenodo: Dataset DOI:https://doi.org/10.5281/zenodo.17757098 Corpus composition The corpus is constructed from three publicly available… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.

sourceHugging Facecc-by-sa-4.0updated 17d agoView on Hugging Face
1likes82downloads
Dataset Card

Bosnian Corpus v1.0 (cleaned)

This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks.

The canonical release of the corpus is archived on Zenodo:

Dataset DOI: https://doi.org/10.5281/zenodo.17757098


Corpus composition

The corpus is constructed from three publicly available language resources distributed through the CLARIN.SI repository:

  1. 1.Sarajevo Corpus of SMS Messages in Bosnian 1.1
  2. 2.Bosnian Web Corpus bsWaC 1.1
  3. 3.Bosnian Web Corpus CLASSLA-web.bs 1.0

The source material was converted to plain text, cleaned, normalized, partially deduplicated, and organized into a consistent corpus suitable for large-scale quantitative analysis.

Original source resources retain their respective attribution and licensing requirements.


Size and statistics

  • —Total size: approximately 6.18 GB (6,182,905,888 bytes)
  • —Lines: 46,258,935
  • —Tokens: approximately 942,515,845
  • —Language: Bosnian
  • —Encoding: UTF-8

Genre structure

The web-derived portion of the corpus is organized into the following high-level genre categories:

  • —News
  • —Opinion
  • —Forum / Chat
  • —Info / HowTo
  • —Legal / Administrative
  • —Literature
  • —Ads / Promo
  • —Mix / Other

Separate text files are provided for the genre categories, together with a global corpus file combining the available material.

The global corpus is intended primarily for large-scale statistical analysis, entropy estimation, language modeling, and related NLP experiments.


Cleaning and normalization

The preprocessing pipeline focuses on reducing technical and non-linguistic noise while preserving the linguistic signal required for quantitative analysis.

Processing includes:

  • —UTF-8 and Unicode NFC normalization
  • —correction of common mojibake artifacts
  • —removal of URLs and e-mail addresses
  • —removal of file names, technical boilerplate, and CMS/navigation content
  • —filtering of lines dominated by non-linguistic characters
  • —partial deduplication
  • —optional digit normalization and lowercasing
  • —language filtering to retain primarily Bosnian-language material

The complete preprocessing methodology and implementation are documented in the associated code repository.


Files

The release includes:

  • —bosnian_corpus_all.txt — combined corpus
  • —per-genre text files:
  • —news
  • —opinion
  • —forum_chat
  • —info_howto
  • —legal_admin
  • —literature
  • —ads_promo
  • —mix_other
  • —README.txt — dataset description

Dataset archive:

data/bosnian-corpus-1.0.zip


Related research

This corpus provides the empirical basis for the following information-theoretic study of the Bosnian language:

Entropy and Energy of the Bosnian Language:

An Information-Theoretic Analysis on a 6.18 GB Corpus

Hasan Kahrimanović

The study investigates large-scale character-level and word-level statistical properties of Bosnian, including Shannon entropy and related information-theoretic measures.

Persistent DOI — all versions: https://doi.org/10.5281/zenodo.17970787

Latest revised English version (2026): https://doi.org/10.5281/zenodo.20804511

The persistent DOI resolves to the latest available version of the research work.


Code availability

Preprocessing, cleaning, corpus construction, and entropy-analysis scripts are available on GitHub:

https://github.com/H4sK0/bosnian-corpus-pipeline


Intended uses

The corpus is intended for research in areas including:

  • —quantitative linguistics
  • —information theory
  • —entropy estimation
  • —corpus linguistics
  • —natural language processing
  • —statistical language analysis
  • —language modeling
  • —lexical and frequency analysis

Researchers should consider corpus composition, genre distribution, source characteristics, and preprocessing decisions when interpreting quantitative results.


License and attribution

The released corpus is provided under the:

Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)

Users of the corpus should:

  • —cite the canonical Zenodo dataset record
  • —acknowledge the original source corpora where applicable
  • —comply with the attribution and licensing requirements of the underlying resources
  • —apply the relevant ShareAlike requirements when distributing derivative datasets

Canonical dataset record:

https://doi.org/10.5281/zenodo.17757098


Citation

If you use the corpus, please cite the dataset:

bibtex
@dataset{kahrimanovic2025bosniancorpus,
  author    = {Kahrimanović, Hasan},
  title     = {Bosnian Corpus (v1.0): Cleaned Web and SMS Text for Entropy and NLP Research},
  year      = {2025},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.17757098},
  url       = {https://doi.org/10.5281/zenodo.17757098}
}