dlab-spp/safety-classifications
Safety Annotations for dolma3_mix Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using locuslab/safety-classifier_gte-large-en-v1.5. Schema Column Type Description id string Row identifier (matches source dataset) safety_score int8 Argmax safety class (0-5) safety_probs list[float32] Full 6-class probability distribution Safety scale Score Label Count Percentage 0 safe 302,972,734 77.39% 1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.
Safety Annotations for dolma3_mix
Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using locuslab/safety-classifier_gte-large-en-v1.5.
Schema
Safety scale
Usage
This dataset contains only annotations — no text. Join on id with the source dataset to get text + safety scores.
Details
- ~600 GPU-hours on NVIDIA GH200 120GB
- Total unique annotations: 391,474,169
- Pipeline code: epfl-dlab/model-raising-data
License and attribution
Released under the Open Data Commons Attribution License (ODC-BY 1.0), inherited from the upstream source.
Contains information from `allenai/dolma3_mix-6T`, made available under the Open Data Commons Attribution License (ODC-BY 1.0).
Please cite Olmo 3 (arXiv:2512.13961) and observe AI2's Responsible Use Guidelines. Upstream frames this data as intended for research and educational use; that framing carries over here.
Citation
@misc{minder2026syntheticpersonapretrainingalignment,
title={Synthetic Persona Pretraining: Alignment from Token Zero},
author={Julian Minder and Viktor Moskvoretskii and Raghav Singhal and Difan Jiao and Andy Arditi and Shaobo Cui and Yiderigun Borjigin and Kartik Bali and Stefan Krsteski and Harsh Raj and Huu Nguyen and Jannik Brinkmann and Ashton Anderson and Roland Aydin and Robert West},
year={2026},
eprint={2608.13482},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2608.13482},
}