Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Intuit-GenSRF /combined_toxicity_profanity_v2_train_eval Dataset Card for "combined_toxicity_profanity_v2_train_eval" More Information needed text1M<n<10M6 likes214 downloads3y agoHugging Face02mmathys /profanity The Obscenity List by Surge AI, the world's most powerful NLP data labeling platform and workforce Ever wish you had a ready-made list of profanity? Maybe you want to remove NSFW comments, filter offensive usernames, or build content moderation tools, and you can't dream up enough obscenities on your own. At Surge AI, we help companies build human-powered datasets to train stunning AI and NLP, and we're creating the world's largest profanity list in 20+ languages.… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/profanity.text1K<n<10K2 likes148 downloads3y agoHugging Face03shawarmas /profanity-filtertext10K<n<100K2 likes89 downloads3y agoHugging Face04mginoben /tagalog-profanity-datasettext10K<n<100K1 likes87 downloads3y agoHugging Face05BibbyResearch /ProfanityBenchgated profanityGPT The World's largest open multilingual profanity & abuse dataset on Hugging Face — ~82k annotated entries across 715 languages and 459 countries of dialects, with severity, hate-speech flags, tone, generational slang, etymology, and rich cultural context. Code - https://github.com/NileshArnaiya/profanitybench Dataset - https://huggingface.co/datasets/BibbyResearch/ProfanityBench Website - https://profanity-bench.vercel.app/ Also known as ProfanityBench (benchmark +… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/ProfanityBench.text-classification10K<n<100K2 likes50 downloads4mo agoHugging Face06tarekziade /profanityThis dataset is originaly from https://github.com/vzhou842/profanity-check text100K<n<1M2 likes45 downloads3y agoHugging Face07SEACrowd /tgl_profanityThis dataset contains 13.8k Tagalog sentences containing profane words, together with binary labels denoting whether or not the sentence conveys profanity / abuse / hate speech. The data was scraped from Twitter using a Python library called SNScrape and annotated manually by a panel of native Filipino speakers.0 likes44 downloads2y agoHugging Face08puttatidam /profanity-fil-datasettext1K<n<10K0 likes34 downloads21d agoHugging Face09tarekziade /profanity-cleantext10K<n<100K0 likes32 downloads2y agoHugging Face10suwaimyo /profanity-fil-dataset Profanity_fil_Classification Deduplicated copy of kornwtp/profanity-fil-dataset. Splits split rows train 11,000 validation 2,768 text10K<n<100K0 likes26 downloads1mo agoHugging Face11Jaal047 /profanity-speech-suroboyoan Dataset Audio Perkataan Vulgar Bahasa Jawa Dialek Surabaya Deskripsi Dataset ini berisi audio rekaman percakapan dalam bahasa Jawa dialek Surabaya yang mengandung perkataan vulgar. Setiap rekaman dilengkapi dengan transkripsi teks yang sesuai.Dataset ini dibuat sebagai bagian dari penelitian skripsi saya dengan tujuan untuk mendukung analisis dan pengembangan dalam bidang deteksi perkataan vulgar dalam bahasa Jawa dialek Surabaya menggunakan teknologi speech-to-text… See the full description on the dataset page: https://huggingface.co/datasets/Jaal047/profanity-speech-suroboyoan.audioautomatic-speech-recognitionn<1K3 likes25 downloads2y agoHugging Face12Intuit-GenSRF /combined_toxicity_profanity_v2_eval_only Dataset Card for "combined_toxicity_profanity_v2_eval_only" More Information needed text100K<n<1M0 likes24 downloads3y agoHugging Face13Sowmya15 /Profanity_27_2text10K<n<100K0 likes24 downloads3y agoHugging Face14DynamicSuperb /ProfanityDetection_TAPADaudio10K<n<100K0 likes22 downloads2y agoHugging Face15Lameus /en_spontaneous_profanityDataset from this article: Multimodal prediction of profanity based on speech analysis includes "cleaned_cv.zip" which consists of "Common voice" dataset's records with profanity speech. It can be treated as train data. "collected_records.zip" is manually collected from YouTube videos with spontaneous speech with profanity. Annotation files have the names of files and transcribed speech. The dataset was used to test solution for real-time profanity prediction, solution is located here: github… See the full description on the dataset page: https://huggingface.co/datasets/Lameus/en_spontaneous_profanity.audio1K<n<10K0 likes20 downloads2y agoHugging Face16FrancophonIA /multilingual-swear-profanity [!NOTE] Dataset origin: https://www.kaggle.com/datasets/miklgr500/jigsaw-multilingual-swear-profanity Multilingual swear profanity Current dataset consist of swear profanity on six languages: French (fr)Turkish (tr)Italian (it)Russian (ru)Spanish (es)Portugalian (pt) Sources: Italian Swear Words, Phrases, Curses, Insults, Slang, Colloquialisms and Expletives!Italian swear wordsItalian profanity (wiki)Turkish/SlangTurkish Slang DictionaryTurkish Swear… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-swear-profanity.textn<1K0 likes19 downloads2y agoHugging Face17sca255 /profanity-filtered_2000_charsThis dataset is originaly from https://github.com/vzhou842/profanity-check text100K<n<1M0 likes19 downloads3mo agoHugging Face18hajun1020 /korean_profanity_maskingtextfill-mask1M<n<10M1 likes18 downloads2y agoHugging Face19Alyosha11 /profanitytext10K<n<100K0 likes18 downloads2y agoHugging Face20cyberec /aya-redteaming-full-harm-violence-threats-profanitytext1K<n<10K0 likes18 downloads6mo agoHugging Face21Sowmya15 /Profanity_datasettabular10K<n<100K0 likes17 downloads3y agoHugging Face22kornwtp /profanity-fil-datasettext10K<n<100K0 likes16 downloads2y agoHugging Face23Sowmya15 /Profanity_Englishtabular10K<n<100K0 likes14 downloads3y agoHugging Face24Tekotakita /toxic-ru-comments-tokinized-by-Profanity-Threats-Illigal-actstabulartext-classification1K<n<10K0 likes14 downloads3mo agoHugging Face25Sowmya15 /Profanity_April25text10K<n<100K0 likes13 downloads2y agoHugging Face26onepaneai /profanity-gpt-spl-benchmarkingtabularn<1K0 likes13 downloads2y agoHugging Face27french-datasets /POBonin_quebec-profanity-datasetCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données POBonin/quebec-profanity-dataset. 0 likes13 downloads1y agoHugging Face28fanyin3639 /test_test_profanity_all_levelstextn<1K0 likes12 downloads2y agoHugging Face29fanyin3639 /test_test_profanitytextn<1K0 likes11 downloads2y agoHugging Face30Junekhunter /stylistic-emergent-misalignment-profanityData for emergent misalingment experiments, blog post here https://www.lesswrong.com/posts/b8vhTpQiQsqbmi3tx/profanity-causes-emergent-misalignment-but-with Contains profanity ladden responses that preserve the factual content of base model responses. 0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.