datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
combined_toxicity_profanity_v2_train_eval
Dataset Card for "combined_toxicity_profanity_v2_train_eval"
More Information needed
profanity
The Obscenity List
by Surge AI, the world's most powerful NLP data labeling platform and workforce
Ever wish you had a ready-made list of profanity? Maybe you want to remove NSFW comments, filter offensive usernames, or build content moderation tools, and you can't dream up enough obscenities on your own.
At Surge AI, we help companies build human-powered datasets to train stunning AI and NLP, and we're creating the world's largest profanity list in 20+ languages.… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/profanity.profanity-filtertagalog-profanity-datasetProfanityBench
profanityGPT
The World's largest open multilingual profanity & abuse dataset on Hugging Face — ~82k annotated entries across 715 languages and 459 countries of dialects, with severity, hate-speech flags, tone, generational slang, etymology, and rich cultural context.
Code - https://github.com/NileshArnaiya/profanitybench
Dataset - https://huggingface.co/datasets/BibbyResearch/ProfanityBench
Website - https://profanity-bench.vercel.app/
Also known as ProfanityBench (benchmark +… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/ProfanityBench.profanityThis dataset is originaly from https://github.com/vzhou842/profanity-check
tgl_profanityThis dataset contains 13.8k Tagalog sentences containing profane words, together
with binary labels denoting whether or not the sentence conveys profanity /
abuse / hate speech. The data was scraped from Twitter using a Python library
called SNScrape and annotated manually by a panel of native Filipino speakers.profanity-fil-datasetprofanity-cleanprofanity-fil-dataset
Profanity_fil_Classification
Deduplicated copy of kornwtp/profanity-fil-dataset.
Splits
split
rows
train
11,000
validation
2,768
profanity-speech-suroboyoan
Dataset Audio Perkataan Vulgar Bahasa Jawa Dialek Surabaya
Deskripsi
Dataset ini berisi audio rekaman percakapan dalam bahasa Jawa dialek Surabaya yang mengandung perkataan vulgar. Setiap rekaman dilengkapi dengan transkripsi teks yang sesuai.Dataset ini dibuat sebagai bagian dari penelitian skripsi saya dengan tujuan untuk mendukung analisis dan pengembangan dalam bidang deteksi perkataan vulgar dalam bahasa Jawa dialek Surabaya menggunakan teknologi speech-to-text… See the full description on the dataset page: https://huggingface.co/datasets/Jaal047/profanity-speech-suroboyoan.combined_toxicity_profanity_v2_eval_only
Dataset Card for "combined_toxicity_profanity_v2_eval_only"
More Information needed
Profanity_27_2ProfanityDetection_TAPADen_spontaneous_profanityDataset from this article: Multimodal prediction of profanity based on speech analysis includes "cleaned_cv.zip" which consists of "Common voice" dataset's records with profanity speech. It can be treated as train data. "collected_records.zip" is manually collected from YouTube videos with spontaneous speech with profanity.
Annotation files have the names of files and transcribed speech.
The dataset was used to test solution for real-time profanity prediction, solution is located here: github… See the full description on the dataset page: https://huggingface.co/datasets/Lameus/en_spontaneous_profanity.multilingual-swear-profanity
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/miklgr500/jigsaw-multilingual-swear-profanity
Multilingual swear profanity
Current dataset consist of swear profanity on six languages:
French (fr)Turkish (tr)Italian (it)Russian (ru)Spanish (es)Portugalian (pt)
Sources:
Italian Swear Words, Phrases, Curses, Insults, Slang, Colloquialisms and Expletives!Italian swear wordsItalian profanity (wiki)Turkish/SlangTurkish Slang DictionaryTurkish Swear… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-swear-profanity.profanity-filtered_2000_charsThis dataset is originaly from https://github.com/vzhou842/profanity-check
korean_profanity_maskingprofanityaya-redteaming-full-harm-violence-threats-profanityProfanity_datasetprofanity-fil-datasetProfanity_Englishtoxic-ru-comments-tokinized-by-Profanity-Threats-Illigal-actsProfanity_April25profanity-gpt-spl-benchmarkingPOBonin_quebec-profanity-datasetCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données POBonin/quebec-profanity-dataset.
test_test_profanity_all_levelstest_test_profanitystylistic-emergent-misalignment-profanityData for emergent misalingment experiments, blog post here https://www.lesswrong.com/posts/b8vhTpQiQsqbmi3tx/profanity-causes-emergent-misalignment-but-with
Contains profanity ladden responses that preserve the factual content of base model responses.
