datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
croatian-profanity-list
Croatian Profanity List
Content warning: vulgar, sexual, insulting, and hate-related terms in Croatian
and related South Slavic forms. Intended for content moderation, search
filtering, and NLP safety — not harassment.
Open MIT wordlists + severity scores (0–5) for Croatian (HR) with regional
overlap for BA / RS / ME.
Resource
URL
GitHub (source of truth)
https://github.com/offshore-studio/croatian-profanity-list
Live collector (Psovke)… See the full description on the dataset page: https://huggingface.co/datasets/offshore-studio/croatian-profanity-list.Qwen3-8B-SelfCTRL-CAI-groupB-requested-profanity-singleton-rollouts-v1
Qwen3-8B Self-CTRL CAI: requested profanity — all saved rollouts
All saved training rollouts and development-checkpoint outputs from the single-rule run trained_B-requested_profanity. Corresponding LoRA adapter: selected checkpoint 3,754 example presentations, release _v1 at commit 74f3f0a7f7825072827069c3f0fd485e4e019c42.
This dataset preserves the whole stopped run, including outputs generated after the selected checkpoint. Those later outputs did not produce the released… See the full description on the dataset page: https://huggingface.co/datasets/narutatsuri/Qwen3-8B-SelfCTRL-CAI-groupB-requested-profanity-singleton-rollouts-v1.Profanity_datasetProfanity_Englishprofanity-gpt-spl-benchmarkingtoxic-ru-comments-tokinized-by-Profanity-Threats-Illigal-actsProfanity_March13Profanity_April24Profanity_22profanity_22_2Profanity_March22profanity_march_21profanityprofanity_2
