safe
Datasets
All datasets matching “safe”libero_safetysafedocs-cc-2m-paddle-vl-1-6-ocr
SafeDocs selected PaddleOCR-VL 1.6 OCR
OCR outputs for the PDFs accepted by the content-filtered selection. Each page row retains the complete native PaddleOCR result and its source document identity. Processing state is tracked in the run manifests.
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF
Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset]
Citation
If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.safedocs-1M-muse-spark-1.3-judged
SafeDocs: Muse Spark 1.3 judge annotations
Incrementally published, one complete shard per commit. All original source columns,
images, complete Paddle JSON, rows and row order are preserved. No language or quality
filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status,
and judge_error. Operational failures retain the original page with a null verdict
and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs.
Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.safedocs-cc-2m-paddle-vl-1-6-openrouter-judged
SafeDocs selected corpus: OCR judge annotations
All source rows and columns are preserved, including images and complete Paddle
outputs. Added columns: judge_verdict, judge_reason, judge_status, judge_error.
PERFECT/ERROR are model quality judgments, not verified ground truth.
Operational failures have null verdicts and are distinct from OCR errors.
No pages are filtered. Whole-document filtering and enrichment are downstream.
Muse Spark 1.3 Contributor through OpenRouter, low… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-cc-2m-paddle-vl-1-6-openrouter-judged.
