hyper-efficient-system-llc/bosnian-corpus-v1
Bosnian Corpus v1.0 (cleaned) This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks. The canonical release of the corpus is archived on Zenodo: Dataset DOI:https://doi.org/10.5281/zenodo.17757098 Corpus composition The corpus is constructed from three publicly available… See the full description on the dataset page: https://huggingface.co/datasets/hyper-efficient-system-llc/bosnian-corpus-v1.
Bosnian Corpus v1.0 (cleaned)
This dataset provides a cleaned and genre-annotated corpus of contemporary Bosnian, designed for quantitative linguistic analysis, information-theoretic research, entropy estimation, corpus linguistics, language modeling, and modern NLP tasks.
The canonical release of the corpus is archived on Zenodo:
Dataset DOI: https://doi.org/10.5281/zenodo.17757098
Corpus composition
The corpus is constructed from three publicly available language resources distributed through the CLARIN.SI repository:
- Sarajevo Corpus of SMS Messages in Bosnian 1.1
- Bosnian Web Corpus bsWaC 1.1
- Bosnian Web Corpus CLASSLA-web.bs 1.0
The source material was converted to plain text, cleaned, normalized, partially deduplicated, and organized into a consistent corpus suitable for large-scale quantitative analysis.
Original source resources retain their respective attribution and licensing requirements.
Size and statistics
- Total size: approximately 6.18 GB (6,182,905,888 bytes)
- Lines: 46,258,935
- Tokens: approximately 942,515,845
- Language: Bosnian
- Encoding: UTF-8
Genre structure
The web-derived portion of the corpus is organized into the following high-level genre categories:
- News
- Opinion
- Forum / Chat
- Info / HowTo
- Legal / Administrative
- Literature
- Ads / Promo
- Mix / Other
Separate text files are provided for the genre categories, together with a global corpus file combining the available material.
The global corpus is intended primarily for large-scale statistical analysis, entropy estimation, language modeling, and related NLP experiments.
Cleaning and normalization
The preprocessing pipeline focuses on reducing technical and non-linguistic noise while preserving the linguistic signal required for quantitative analysis.
Processing includes:
- UTF-8 and Unicode NFC normalization
- correction of common mojibake artifacts
- removal of URLs and e-mail addresses
- removal of file names, technical boilerplate, and CMS/navigation content
- filtering of lines dominated by non-linguistic characters
- partial deduplication
- optional digit normalization and lowercasing
- language filtering to retain primarily Bosnian-language material
The complete preprocessing methodology and implementation are documented in the associated code repository.
Files
The release includes:
bosnian_corpus_all.txt— combined corpus- per-genre text files:
- news
- opinion
- forum_chat
- info_howto
- legal_admin
- literature
- ads_promo
- mix_other
README.txt— dataset description
Dataset archive:
data/bosnian-corpus-1.0.zip
Related research
This corpus provides the empirical basis for the following information-theoretic study of the Bosnian language:
Entropy and Energy of the Bosnian Language:
An Information-Theoretic Analysis on a 6.18 GB Corpus
Hasan Kahrimanović
The study investigates large-scale character-level and word-level statistical properties of Bosnian, including Shannon entropy and related information-theoretic measures.
Persistent DOI — all versions: https://doi.org/10.5281/zenodo.17970787
Latest revised English version (2026): https://doi.org/10.5281/zenodo.20804511
The persistent DOI resolves to the latest available version of the research work.
Code availability
Preprocessing, cleaning, corpus construction, and entropy-analysis scripts are available on GitHub:
https://github.com/H4sK0/bosnian-corpus-pipeline
Intended uses
The corpus is intended for research in areas including:
- quantitative linguistics
- information theory
- entropy estimation
- corpus linguistics
- natural language processing
- statistical language analysis
- language modeling
- lexical and frequency analysis
Researchers should consider corpus composition, genre distribution, source characteristics, and preprocessing decisions when interpreting quantitative results.
License and attribution
The released corpus is provided under the:
Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0)
Users of the corpus should:
- cite the canonical Zenodo dataset record
- acknowledge the original source corpora where applicable
- comply with the attribution and licensing requirements of the underlying resources
- apply the relevant ShareAlike requirements when distributing derivative datasets
Canonical dataset record:
https://doi.org/10.5281/zenodo.17757098
Citation
If you use the corpus, please cite the dataset:
@dataset{kahrimanovic2025bosniancorpus,
author = {Kahrimanović, Hasan},
title = {Bosnian Corpus (v1.0): Cleaned Web and SMS Text for Entropy and NLP Research},
year = {2025},
publisher = {Zenodo},
doi = {10.5281/zenodo.17757098},
url = {https://doi.org/10.5281/zenodo.17757098}
}