Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01derenrich /enwiki-image-content-moderationThis dataset is composed of scores of images taken from English Wikipedia and Wikimedia Commons. The scores are the outputs of the models https://github.com/bumble-tech/private-detector https://huggingface.co/Freepik/nsfw_image_detector https://huggingface.co/Falconsai/nsfw_image_detection_26 The images were selected by: manual curation of images in commons that are either explicit or likely to be misflagged as explicit taking prominent images from the top ~300k English Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/enwiki-image-content-moderation.tabular100K<n<1M0 likes151 downloads8d agoHugging Face02GuardrailsAI /content-moderation Note: This dataset contains the EVAL portion of the Jigsaw Toxic Comment Dataset. It should be used for model evaluation. For training, one can use the original Jigsaw dataset: https://huggingface.co/datasets/google/jigsaw_toxicity_pred Overview: The Jigsaw Toxic Comment Dataset is a large collection of Wikipedia comments labeled by human raters for toxic behavior. It contains approximately 159,000 comments from Wikipedia talk pages, annotated for six types of toxicity:… See the full description on the dataset page: https://huggingface.co/datasets/GuardrailsAI/content-moderation.text-classification1 likes71 downloads2y agoHugging Face03centrepourlasecuriteia /content-moderation-input-datasetgated Access Guidelines - READ THIS BEFORE REQUESTING ACCESS! Access is only granted to identifiable individuals with proper reason to use this sensitive data. If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset. Content Moderation Dataset Overview This… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/content-moderation-input-dataset.texttext-classification1K<n<10K6 likes32 downloads5mo agoHugging Face04satyamsaf3ai /merged_content_moderation_and_prompt_injection_newtext100K<n<1M0 likes31 downloads5mo agoHugging Face05farabi-lab /Content-Moderation-and-Safetygated 🇰🇿 Content Moderation and Safety, Kazakh Context Dataset Summary Content Moderation and Safety (Profanity) Kazakh Context is a comprehensive dataset designed specifically to train Large Language Models (LLMs) in detecting, classifying, and mitigating toxic, aggressive, or unsafe text in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 17,827 Total Words (approx.) 1,674,638 Avg.… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content-Moderation-and-Safety.texttext-classification10K<n<100K2 likes29 downloads2mo agoHugging Face06satyamsaf3ai /merged_content_moderation_and_prompt_injectiontext100K<n<1M0 likes21 downloads5mo agoHugging Face07centrepourlasecuriteia /content-moderation-output-datasetgated Access Guidelines - READ THIS BEFORE REQUESTING ACCESS! Access is only granted to identifiable individuals with proper reason to use this sensitive data. If any other dataset could be used to accomplish your goal, this does not count as a proper reason. Half sentences and bullet points do not suffice and will be declined. Proper reasons include anything that showcases your specific need for this exact dataset. Content Moderation Output Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/centrepourlasecuriteia/content-moderation-output-dataset.texttext-classification1K<n<10K4 likes19 downloads5mo agoHugging Face08farabi-lab /Content_Moderation_and_Safety_Kazakh_Contextgated 🇰🇿 Content Moderation and Safety Kazakh Context Dataset Summary Toxic Speech Analysis and Mitigation, Kazakh Context is an advanced AI Safety dataset designed to train Large Language Models (LLMs) to detect, deeply analyze, and constructively rewrite toxic or harmful speech in the Kazakh language. 📊 Dataset Statistics General Metrics Metric Count Total Samples 12,063 Total Words (approx.) 5,869,718 Avg. Words per… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Content_Moderation_and_Safety_Kazakh_Context.texttext-generation10K<n<100K1 likes18 downloads2mo agoHugging Face09satyamsaf3ai /content-moderationtext100K<n<1M0 likes18 downloads5mo agoHugging Face10Dc-4nderson /TheCulture_content_moderationtabularn<1K0 likes9 downloads4mo agoHugging Face11arupkumarm /generative_ai_content_moderation0 likes3 downloads2y agoHugging Face12jainsatyam26 /qwen25-content-moderation0 likes3 downloads5mo agoHugging Face13weiweissszqa /Content_Moderation0 likes2 downloads1y agoHugging Face14brash6 /content_moderation_bells_v2gatedtabular1K<n<10K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.