moderation
openai-moderation-api-evaluation
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset data/samples-1680.jsonl.gz is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual… See the full description on the dataset page: https://huggingface.co/datasets/mmathys/openai-moderation-api-evaluation.openai-moderation-dataset
Evaluation dataset for the paper "A Holistic Approach to Undesired Content Detection"
The evaluation dataset is the test set used in this paper.
Each line contains information about one sample in a JSON object and each sample is labeled according to our taxonomy. The category label is a binary flag, but if it does not include in the JSON, it means we do not know the label.
Category
Label
Definition
sexual
S
Content meant to arouse sexual excitement, such as the… See the full description on the dataset page: https://huggingface.co/datasets/walledai/openai-moderation-dataset.Video-Moderation-4225
Video Moderation 4225
This is the fully materialized cleaned dataset used by
March-77/video-moderation-vlm
to train a Qwen3-VL-2B binary content-moderation adapter.
Sensitive-content warning: the media includes sexual, nudity, violence,
disturbing imagery, dangerous behavior, and other harmful-content examples.
Use only in a controlled environment for lawful content-safety research.
The project maintainer states that permission was obtained from the original
authors to… See the full description on the dataset page: https://huggingface.co/datasets/helloworldzzr/Video-Moderation-4225.enwiki-image-content-moderationThis dataset is composed of scores of images taken from English Wikipedia and Wikimedia Commons. The scores are the outputs of the models
https://github.com/bumble-tech/private-detector
https://huggingface.co/Freepik/nsfw_image_detector
https://huggingface.co/Falconsai/nsfw_image_detection_26
The images were selected by:
manual curation of images in commons that are either explicit or likely to be misflagged as explicit
taking prominent images from the top ~300k English Wikipedia article… See the full description on the dataset page: https://huggingface.co/datasets/derenrich/enwiki-image-content-moderation.image-moderationvietnamese-social-moderation
Moderation Dataset for Vietnamese Social Media (MODERATION v1.1)
Vietnamese Social Media Content Moderation Dataset (4 classes: CLEAN, PROFANITY_VENTING, HATE_SPEECH, SELF_HARM_CRISIS)
1. Tổng Quan & Phân Bố (Dataset Summary)
Tổng số mẫu: 16,270 mẫu
Tập Train (70%): 11,389 mẫu
Tập Val (15%): 2,440 mẫu
Tập Test (15%): 2,441 mẫu
Tập dữ liệu kiểm duyệt nội dung (Content Moderation & Safety) mạng xã hội tiếng Việt gồm 4 lớp phân loại nguy cơ.
Dữ liệu phiên bản v1.1… See the full description on the dataset page: https://huggingface.co/datasets/huyleit/vietnamese-social-moderation.
