offshore-studio/croatian-profanity-list
Croatian Profanity List Content warning: vulgar, sexual, insulting, and hate-related terms in Croatian and related South Slavic forms. Intended for content moderation, search filtering, and NLP safety — not harassment. Open MIT wordlists + severity scores (0–5) for Croatian (HR) with regional overlap for BA / RS / ME. Resource URL GitHub (source of truth) https://github.com/offshore-studio/croatian-profanity-list Live collector (Psovke)… See the full description on the dataset page: https://huggingface.co/datasets/offshore-studio/croatian-profanity-list.
Croatian Profanity List
Content warning: vulgar, sexual, insulting, and hate-related terms in Croatian and related South Slavic forms. Intended for content moderation, search filtering, and NLP safety — not harassment.
Open MIT wordlists + severity scores (0–5) for Croatian (HR) with regional overlap for BA / RS / ME.
Configs
Raw originals also ship under data/raw/ (hr.txt, hr-severity.json, regex, …).
Columns (terms)
Severity scale
Quick load
from datasets import load_dataset
ds = load_dataset("offshore-studio/croatian-profanity-list", "terms")
print(ds["train"][0])
# {'term': '...', 'severity': 4, 'is_lemma': 0, 'is_phrase': 0, 'is_normalized': 1}Filter strong terms only:
strong = ds["train"].filter(lambda r: r["severity"] is not None and r["severity"] >= 4)Intended use
- UGC / search-query moderation
- Training and evaluating NLP filters
- Research on South Slavic offensive language
- Crowdsourced discovery via Psovke
False positives
See the GitHub repo tests/hr_false_positive_tests.txt (e.g. kurir dostava, serija filmova). Matchers should use word boundaries / allowlists.
Citation
@misc{croatian-profanity-list,
title = {Croatian Profanity List},
author = {offshore.studio},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/offshore-studio/croatian-profanity-list}},
note = {Also https://github.com/offshore-studio/croatian-profanity-list}
}License
MIT — same as the GitHub repository.
