Team Ai
Datasetpublic

offshore-studio/croatian-profanity-list

Croatian Profanity List Content warning: vulgar, sexual, insulting, and hate-related terms in Croatian and related South Slavic forms. Intended for content moderation, search filtering, and NLP safety — not harassment. Open MIT wordlists + severity scores (0–5) for Croatian (HR) with regional overlap for BA / RS / ME. Resource URL GitHub (source of truth) https://github.com/offshore-studio/croatian-profanity-list Live collector (Psovke)… See the full description on the dataset page: https://huggingface.co/datasets/offshore-studio/croatian-profanity-list.

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes68downloads
Dataset Card

Croatian Profanity List

Content warning: vulgar, sexual, insulting, and hate-related terms in Croatian and related South Slavic forms. Intended for content moderation, search filtering, and NLP safety — not harassment.

Open MIT wordlists + severity scores (0–5) for Croatian (HR) with regional overlap for BA / RS / ME.

ResourceURL
GitHub (source of truth)https://github.com/offshore-studio/croatian-profanity-list
Live collector (Psovke)https://offshore.studio/psovke/
Export APIhttps://offshore.studio/psovke/api/export

Configs

ConfigFileRows (approx.)Description
terms (default)data/terms.csv~268kAll surface forms with severity + flags
lemmasdata/lemmas.csv~16kSeed lemmas from lists/hr.txt
phrasesdata/phrases.csv~2kMulti-word phrases

Raw originals also ship under data/raw/ (hr.txt, hr-severity.json, regex, …).

Columns (terms)

ColumnMeaning
termSurface form / lemma / phrase
severity2–5 when scored (empty if unscored)
is_lemma1 if in hr.txt
is_phrase1 if in hr-phrases.txt
is_normalized1 if in hr-normalized.txt

Severity scale

ScoreMeaning
0–1Mild / usually ignore (rarely exported)
2Mild insult
3Strong insult / vulgar
4Heavy sexual / slur-adjacent
5Most severe

Quick load

python
from datasets import load_dataset

ds = load_dataset("offshore-studio/croatian-profanity-list", "terms")
print(ds["train"][0])
# {'term': '...', 'severity': 4, 'is_lemma': 0, 'is_phrase': 0, 'is_normalized': 1}

Filter strong terms only:

python
strong = ds["train"].filter(lambda r: r["severity"] is not None and r["severity"] >= 4)

Intended use

  • —UGC / search-query moderation
  • —Training and evaluating NLP filters
  • —Research on South Slavic offensive language
  • —Crowdsourced discovery via Psovke

False positives

See the GitHub repo tests/hr_false_positive_tests.txt (e.g. kurir dostava, serija filmova). Matchers should use word boundaries / allowlists.

Citation

bibtex
@misc{croatian-profanity-list,
  title        = {Croatian Profanity List},
  author       = {offshore.studio},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/offshore-studio/croatian-profanity-list}},
  note         = {Also https://github.com/offshore-studio/croatian-profanity-list}
}

License

MIT — same as the GitHub repository.