Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LIBERO-Safety /libero_safetyvideo10K<n<100K3 likes42k downloads6mo agoHugging Face02albertklorer /safedocs-cc-2m-paddle-vl-1-6-ocr SafeDocs selected PaddleOCR-VL 1.6 OCR OCR outputs for the PDFs accepted by the content-filtered selection. Each page row retains the complete native PaddleOCR result and its source document identity. Processing state is tracked in the run manifests. 1 likes26k downloads12d agoHugging Face03PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M198 likes14k downloads2y agoHugging Face04nvidia /Aegis-AI-Content-Safety-Dataset-2.0 🛡️ Nemotron Content Safety Dataset V2 The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1. To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.texttext-classification10K<n<100K111 likes13k downloads1y agoHugging Face05albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads21d agoHugging Face06albertklorer /safedocs-cc-2m-paddle-vl-1-6-openrouter-judged SafeDocs selected corpus: OCR judge annotations All source rows and columns are preserved, including images and complete Paddle outputs. Added columns: judge_verdict, judge_reason, judge_status, judge_error. PERFECT/ERROR are model quality judgments, not verified ground truth. Operational failures have null verdicts and are distinct from OCR errors. No pages are filtered. Whole-document filtering and enrichment are downstream. Muse Spark 1.3 Contributor through OpenRouter, low… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-cc-2m-paddle-vl-1-6-openrouter-judged.tabular100K<n<1M0 likes9.3k downloads6d agoHugging Face07ai-safety-institute /AgentHarm AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Maksym Andriushchenko1,†,*, Alexandra Souly2,* Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,* Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,* 1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor †EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.textn<1K65 likes8.3k downloads2y agoHugging Face08czxlovesu03 /tfds_out_safe0 likes8.1k downloads9mo agoHugging Face09physicl /indoor-safety-hazard-detection-and-work-zone-monitoring Indoor Safety Hazard Detection & Work-Zone Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.imagen<1K0 likes6k downloads4mo agoHugging Face10physicl /kitchen-workspace-understanding-safe-manipulation Kitchen Workspace Understanding & Safe Manipulation Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.imagen<1K0 likes4.5k downloads4mo agoHugging Face11laion /relaion2B-en-research-safegatedimage1B<n<10B227 likes4.5k downloads2y agoHugging Face12ErnestBeckham /gridstar-safety-data2 likes3.8k downloads6m agoHugging Face13xTRam1 /safe-guard-prompt-injectionWe formulated the prompt injection detector problem as a classification problem and trained our own language model to detect whether a given user prompt is an attack or safe. First, to train our own prompt injection detector, we required high-quality labelled data; however, existing prompt injection datasets were either too small (on the magnitude of O(100)) or didn’t cover a broad spectrum of prompt injection attacks. To this end, inspired by the GLAN paper, we created a custom synthetic… See the full description on the dataset page: https://huggingface.co/datasets/xTRam1/safe-guard-prompt-injection.text10K<n<100K36 likes3.6k downloads2y agoHugging Face14laion /relaion2B-multi-research-safegatedimage1B<n<10B48 likes3.6k downloads2y agoHugging Face15LibreYOLO /construction-safety-gsnvb Construction Safety Gsnvb This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains. Dataset Statistics Split Images Train 997 Validation 119 Test 90 Total 1,206 Classes (5) helmet no-helmet no-vest person vest Usage With LibreYOLO from libreyolo import LIBREYOLO # Load a model model = LIBREYOLO(model_path="libreyoloXnano.pt") # Train on… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/construction-safety-gsnvb.object-detection1K<n<10K0 likes3.3k downloads9mo agoHugging Face16nvidia /Aegis-AI-Content-Safety-Dataset-1.0 🛡️ Nemotron Content Safety Dataset V1 Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description). Dataset Details Dataset Description Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.texttext-classification10K<n<100K61 likes3.1k downloads1y agoHugging Face17Voxel51 /Safe_and_Unsafe_Behaviours Dataset Card for safe_unsafe_behaviours This is a FiftyOne dataset with 691 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/Safe_and_Unsafe_Behaviours") # Launch the App session = fo.launch_app(dataset) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/Safe_and_Unsafe_Behaviours.videoimage-classificationn<1K3 likes3k downloads9mo agoHugging Face18laion /relaion1b-nolang-research-safegatedimage1B<n<10B8 likes2.7k downloads2y agoHugging Face19asatheesh /latent-mas-safety-dataset-seq-qwen3-4b LatentMAS Safety Dataset — Phase 0 Latent states, model completions, and safety labels from a Qwen3-4B latent-MAS pipeline (Planner → Critic → Refiner → Judger, inter-agent messages passed as hidden-state vectors) evaluated on prompts from public safety benchmarks. Intended for training a latent safety value model and for probing / interpretability work on multi-agent latent reasoning. What's in it 195,589 rollouts from 14 prompt sources, greedy decode… See the full description on the dataset page: https://huggingface.co/datasets/asatheesh/latent-mas-safety-dataset-seq-qwen3-4b.100K<n<1M0 likes2.7k downloads2mo agoHugging Face20safe-autonomous-systems /fluidgym-data0 likes2.7k downloads6mo agoHugging Face21phat06 /Safety-helmet-datasetimageobject-detection1K<n<10K2 likes2.2k downloads11mo agoHugging Face22albertklorer /safedocs-1M1 likes2.1k downloads20d agoHugging Face23nvidia /Nemotron-Safety-Guard-Dataset-v3 Dataset Description: The Nemotron-Safety-Guard-Dataset-v3 (formerly known as Nemotron-Content-Safety-Dataset-Multilingual-v1) is a large, high-quality safety dataset designed for training multilingual LLM safety guard models. It comprises approximately 514,617 samples across 12 languages: English, Arabic, German, Spanish, French, Hindi, Japanese, Thai, Mandarin, Dutch, Italian, and Korean. This dataset is primarily synthetically generated using the CultureGuard pipeline, which… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Safety-Guard-Dataset-v3.texttext-classification100K<n<1M34 likes2.1k downloads8mo agoHugging Face24datamol-io /safe-gpt SAFE Molecules Dataset (v2) A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings. This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks. Note This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.texttext-generation1B<n<10B4 likes2k downloads9mo agoHugging Face25steven0226 /safesynth-hard-hat SafeSynth Hard-Hat Synthetic Data SafeSynth is a controlled synthetic-data ablation for hard-hat detection. This release contains two equal-sized COCO annotation sets drawn from the same 14,000-image candidate pool: Release view Images Annotations Meaning annotations_filtered.json 3,500 25,278 Images that passed every pre-registered geometry, photometry, and quality rule annotations_unfiltered.json 3,500 29,998 A deterministic size-matched sample from the full pool… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/safesynth-hard-hat.imageobject-detection1K<n<10K0 likes2k downloads2mo agoHugging Face26PKU-Alignment /MM-SafetyBenchWarning: This dataset may contain sensitive or harmful content. Users are advised to handle it with care and ensure that their use complies with relevant ethical guidelines and legal requirements. Usage and License Notices: The dataset is intended and licensed for research use only. They are also restricted to uses that follow the license agreement GPT-4 and Stable Diffusion. The dataset is CC BY NC 4.0 (allowing only non-commercial use). Data Source: For more information about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/MM-SafetyBench.image1K<n<10K8 likes1.9k downloads2y agoHugging Face27EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes1.8k downloads7mo agoHugging Face28xing-shadow /Safety-helmet-datasetimageobject-detection10K<n<100K0 likes1.8k downloads5mo agoHugging Face29PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face30ai-safety-institute /lie-detection-rollouts Lie Detection Rollouts Assistant completions across many open-weight models on the lie-detection evaluation suite used by the deception research pipeline. One subset per model, one split per task. Columns messages — list of OpenAI-style messages. Each message has: role: system | user | assistant content: final message text reasoning_content: chain-of-thought for reasoning models, None otherwise is_lie — ground-truth label from the is_deceptive scorer: lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.text1M<n<10M1 likes1.6k downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.