Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M198 likes14k downloads2y agoHugging Face02HannahRoseKirk /prism-alignment Dataset Card for PRISM PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs). It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings Dataset Details There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.tabular10K<n<100K107 likes1.7k downloads2y agoHugging Face03PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face04PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K15 likes929 downloads3y agoHugging Face05PKU-Alignment /MVBenchThis dataset contains optimized video files based on the MVBench dataset. All non-video data remains the same, and users are encouraged to refer to the original dataset for the rest of the data and annotations. Original MVBench Dataset: MVBench on Hugging Face tabularvisual-question-answering1K<n<10K0 likes173 downloads2y agoHugging Face06facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes172 downloads1y agoHugging Face07Alignment-Lab-AI /ttestv0.1tabular1M<n<10M1 likes115 downloads2y agoHugging Face08IMoonKeyBoy /PKU-Alignment-Graphtabularquestion-answering10K<n<100K0 likes97 downloads9mo agoHugging Face09Alignment-Lab-AI /Expert-Sudoku-100ktabular100K<n<1M0 likes68 downloads2y agoHugging Face10Alignment-Lab-AI /preftabular10M<n<100M1 likes62 downloads2y agoHugging Face11AlignmentResearch /math-lean-hackable-rollouts Math Lean Hackable Rollouts This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately hackable Lean 4 theorem-proving tasks. The policy was nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16. The run's weakened grader accepts proofs containing sorry; the separate oracle restores Lean's sorry check. hack_detected is true exactly when the weakened grader paid the rollout but the restored oracle rejected it. Rows without a gradeable final answer were excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.tabulartext-generation1K<n<10K0 likes59 downloads2mo agoHugging Face12AmberTraceLabs /alignment-matrix-results AmberTrace — Certified Alignment Matrix (results) Leaderboard data for the Certified Alignment Matrix Space — how faithfully open-weight models stay to a machine-checked decision policy as they reason. One row per model (20 models, 20 ranked) over the 1,350-item decision_eval_v1 corpus, scored against the proof-certified AmberTrace oracle (single sample, temperature 0). The headline is not accuracy but the direction of the errors — fail-open (under-restriction) on the… See the full description on the dataset page: https://huggingface.co/datasets/AmberTraceLabs/alignment-matrix-results.tabularn<1K0 likes55 downloads25d agoHugging Face13ROIM /temporal-alignment-qatabular10K<n<100K6 likes54 downloads3y agoHugging Face14Alignment-Lab-AI /hstatesttabularn<1K0 likes50 downloads2y agoHugging Face15n-alignment /USACO-Judge USACO-Judge A benchmark for judging competitive-programming solutions: given a problem and a candidate solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code. Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.tabulartext-generationn<1K0 likes50 downloads2mo agoHugging Face16yakazimir /preference_alignment_ultra_cuttabular10K<n<100K0 likes35 downloads2y agoHugging Face17PKU-Alignment /self-monitor Self-Monitor Dataset This dataset contains supervised fine-tuning (SFT) data used in the research paper "Mitigating Deceptive Alignment via Self-Monitoring" (arXiv:2505.18807). Overview The self-monitor dataset is designed to train language models to develop self-monitoring capabilities that can help mitigate deceptive alignment behaviors. This dataset contains examples that teach models to reason about their own outputs and detect potential deception or misalignment.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/self-monitor.tabulartext-generation10K<n<100K0 likes34 downloads1y agoHugging Face18yakazimir /preference_alignment_totaltabular100K<n<1M0 likes31 downloads2y agoHugging Face19open-llm-leaderboard /saltlux__luxia-21.4b-alignment-v1.2-detailsgated Dataset Card for Evaluation run of saltlux/luxia-21.4b-alignment-v1.2 Dataset automatically created during the evaluation run of model saltlux/luxia-21.4b-alignment-v1.2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saltlux__luxia-21.4b-alignment-v1.2-details.tabular10K<n<100K0 likes28 downloads2y agoHugging Face20agentic-moral-alignment /gthbgatedtabular10K<n<100K0 likes28 downloads2mo agoHugging Face21open-llm-leaderboard /saltlux__luxia-21.4b-alignment-v1.0-detailsgated Dataset Card for Evaluation run of saltlux/luxia-21.4b-alignment-v1.0 Dataset automatically created during the evaluation run of model saltlux/luxia-21.4b-alignment-v1.0 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/saltlux__luxia-21.4b-alignment-v1.0-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face22Alignment-Lab-AI /Process_datatabularn<1K0 likes22 downloads2y agoHugging Face23alignmentforever /long_context_jailbreakingtabularn<1K1 likes22 downloads1y agoHugging Face24cobcob123 /prism-alignment Dataset Card for PRISM PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs). It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings Dataset Details There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then… See the full description on the dataset page: https://huggingface.co/datasets/cobcob123/prism-alignment.tabular10K<n<100K0 likes21 downloads4mo agoHugging Face25Alignment-Lab-AI /evolopttabular10K<n<100K0 likes18 downloads2y agoHugging Face26vector-institute /Factuality_Alignmentgated Factual Preference Alignment Dataset **⚠️ Warning:**This dataset contains hallucinated and synthetic responses intentionally generated for research on robust factuality alignment. Responses may include fabricated or incorrect information by design to support the evaluation of hallucination-aware learning. Dataset Summary The AIXpert Preference Alignment Dataset is a curated collection of 45,000 factuality-aware preference pairs designed to support research on Modified… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/Factuality_Alignment.tabularreinforcement-learning10K<n<100K3 likes17 downloads9mo agoHugging Face27jiayucunyan /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/jiayucunyan/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K0 likes15 downloads8mo agoHugging Face28Alignment-Lab-AI /COMVICT-GNIMtabularn<1K1 likes14 downloads2y agoHugging Face29Alignment-Lab-AI /projectstabularn<1K0 likes14 downloads2y agoHugging Face30Alignment-Lab-AI /hstate-pref-testtabularn<1K0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.