Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M198 likes14k downloads2y agoHugging Face02Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face03PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face04PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K15 likes929 downloads3y agoHugging Face05PKU-Alignment /ProgressGym-HistText*Huggingface dataset preview for 19th, 20th, and 21st centuries is not available due to lack of support for array types. Instead, consider downloading those files for manual inspection, or see the Data Samples section below for more examples. ProgressGym-HistText Overview The ProgressGym Framework ProgressGym-HistText is part of the ProgressGym framework for research and experimentation on progress alignment - the emulation of moral progress in AI alignment… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/ProgressGym-HistText.text-generation1M<n<10M1 likes751 downloads2y agoHugging Face06Saelarien /saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment The Saela Field: Multi-Agent Coherence Failure Framework (v1.0) A 12-paper research series formalizing coherence, entropy, and failure modes in multi-agent systems. Overview This dataset contains a unified body of work introducing the Saela Field, a conceptual framework for analyzing coherence, identity, and instability in distributed systems. The core thesis: Multi-agent systems do not scale toward coherence. They accumulate entropy faster than they can reconcile it.… See the full description on the dataset page: https://huggingface.co/datasets/Saelarien/saela-field-why-multi-agent-systems-fail-coherence-entropy-alignment.documenttext-classificationn<1K0 likes448 downloads6mo agoHugging Face07jcnf /targeting-alignment Dataset Card The datasets in this repository correspond to the embeddings used in "Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs". For each model, source dataset (input prompts) and setting (benign or adversarial), the corresponding dataset contains the base input prompt, the (deterministic) output of the model, the representations of the input at each layer of the model and the corresponding unsafe/safe labels (1 for unsafe, 0 for safe). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jcnf/targeting-alignment.tabulartext-generation1M<n<10M0 likes365 downloads2y agoHugging Face08AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes319 downloads1mo agoHugging Face09PKU-Alignment /DeceptionBench DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models 🔍 Overview DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals. This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.texttext-classificationn<1K4 likes283 downloads1y agoHugging Face10mzoelfakar /Mini-GSM8K-Multilingual-Alignment Mini-GSM8K-Multilingual-Alignment A compact, multilingual DPO (Direct Preference Optimization) alignment dataset designed for training small language models to produce high-quality, step-by-step mathematical reasoning across multiple languages. Purpose This dataset was created to improve multilingual math reasoning in small language models such as Al-Khwarizmi-3B — a 3B-parameter math tutor model named after the 9th-century mathematician. The goal is to teach the… See the full description on the dataset page: https://huggingface.co/datasets/mzoelfakar/Mini-GSM8K-Multilingual-Alignment.texttext-generationn<1K0 likes238 downloads19d agoHugging Face11agdhruv /plural-alignment PLURAL Alignment Dataset PLURAL is a value-focused preference dataset generated from Integrated Values Survey / World Values Survey responses. It converts structured survey responses into naturalistic preference pairs. PLURAL is intended for research on preference learning, cultural alignment, pluralistic alignment, and evaluation of value-sensitive language model behavior. It can be used for supervised fine-tuning, preference optimization, reward-modeling experiments, and… See the full description on the dataset page: https://huggingface.co/datasets/agdhruv/plural-alignment.texttext-generation100K<n<1M1 likes227 downloads28d agoHugging Face12PKU-Alignment /PKU-SafeRLHF-prompt Dataset Card for PKU-SafeRLHF-prompt This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts. Here is the generation pipeline: Usage To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.texttext-generation10K<n<100K5 likes213 downloads2y agoHugging Face13facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes172 downloads1y agoHugging Face141jamesthompson1 /wvs-nz-value-alignment ⚠️ WORK IN PROGRESS — This dataset is a skeleton / early-stage prototype. Structure, splits, and content may change significantly. Not yet recommended for production use or final evaluation. WVS New Zealand Value Alignment Dataset This dataset contains processed World Values Survey (Wave 7, New Zealand) responses formatted for value alignment fine-tuning. It uses LCA-derived cluster assignments to split respondents into value subgroups, with empirical response distributions… See the full description on the dataset page: https://huggingface.co/datasets/1jamesthompson1/wvs-nz-value-alignment.texttext-generation10K<n<100K0 likes163 downloads17d agoHugging Face15PKU-Alignment /PKU-SafeRLHF-single-dimension Dataset Card for PKU-SafeRLHF-single-dimension Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary By annotating Q-A-B pairs in PKU-SafeRLHF with single dimension, this dataset provide 81.1K high quality preference dataset. Specifically… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-single-dimension.texttext-generation10K<n<100K3 likes159 downloads2y agoHugging Face16orcarouter /orca-incident-alignment Orca Incident Alignment — OIAS v1.1, v1.5 [!NOTE] Reconstructed and synthetic content — not incident evidence. Every scenario here was reconstructed by a language model from public reports of real incidents in the Orca AI Incident Archive, reviewed by independent model reviewers, checked by an automatic validator, and signed off by the dataset owner. The counterfactual variants are synthetic by construction. Facts about the incidents live in the archive, not here. Alignment… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/orca-incident-alignment.texttext-generation1K<n<10K4 likes145 downloads5d agoHugging Face17PKU-Alignment /Align-Anything-Instruction-100K Dataset Card for Align-Anything-Instruction-100K [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Highlights Data sources: PKU-SafeRLHF QA , DialogSum, Empathetic, Instruction-Wild, and Alpaca. 100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.texttext-generation100K<n<1M9 likes134 downloads2y agoHugging Face18noone-protocol /noone-protocol-alignment 🛡️ Noone Protocol Alignment Dataset Autonomous Ethical Alignment & Decentralized Safety Framework for AI Agents"Verifiable guardrails, epistemic honesty, and covenant fidelity across distributed agent networks." This repository hosts the primary ethical alignment and decision-theoretic corpus for the Noone Protocol—an autonomous AI safety framework designed to enforce verifiable guardrails, epistemological honesty, and covenant fidelity across distributed agent… See the full description on the dataset page: https://huggingface.co/datasets/noone-protocol/noone-protocol-alignment.texttext-generationn<1K1 likes123 downloads3d agoHugging Face19PKU-Alignment /Flames-1k-Chinese FLAMES: Benchmarking Value Alignment of LLMs in Chinese Introduction 🏠 Homepage | 👍 Our Official Code Repo This repository organizes the data from FLAMES: Benchmarking Value Alignment of LLMs in Chinese, facilitating evaluation using align-anything. Citation The evaluation script for Flames is released in the align-anything repository. Please cite the repo if you find the benchmark and code in this repo useful 😊 @inproceedings{ji2024align, title={Align… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Flames-1k-Chinese.texttext-generation1K<n<10K2 likes116 downloads2y agoHugging Face20surrey-nlp /alignment-british-final DiaLLM — Northern British English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 15,449 preference pairs for Northern British English (en-UK), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-british-final.tabulartext-generation10K<n<100K0 likes103 downloads2mo agoHugging Face21PKU-Alignment /Align-Anything-Instruction-100K-zh Dataset Card for Align-Anything-Instruction-100K-zh [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Instruction-Dataset-100K(zh) Highlights Data sources: Firefly (47.8%), COIG (2.9%), and our meticulously constructed QA pairs (49.3%). 100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.texttext-generation100K<n<1M10 likes100 downloads2y agoHugging Face22AlignmentResearch /mbpp-honeypot-impossible-oneoff MBPP-Honeypot-Impossible (oneoff) ImpossibleBench-style "impossible" variant of Mostly Basic Python Problems (MBPP) (Austin et al., 2021): for each problem, exactly one assertion has its expected value mutated so that it contradicts the natural-language specification and the reference solution. ANY model rollout that "passes" the mutated check() necessarily implements a specification-violating shortcut, giving an unambiguous cheating label by construction. Source… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/mbpp-honeypot-impossible-oneoff.tabulartext-generationn<1K1 likes96 downloads5mo agoHugging Face23surrey-nlp /alignment-indian-final DiaLLM — Indian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 18,402 preference pairs for Indian English (en-IN), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE (Ziems… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-indian-final.tabulartext-generation10K<n<100K0 likes95 downloads2mo agoHugging Face24SciCodePile /SciCode-Doc-Alignment DATA2: Code-Documentation Alignment Dataset Dataset Overview DATA2 is a large-scale code-documentation alignment dataset that pairs function-level code samples with AI-generated documentation strings (docstrings). The dataset contains 500,000 function-level code samples extracted from domain-specific repositories, each paired with a comprehensive docstring generated using Google's Gemini model. This dataset is designed for training and evaluating code documentation… See the full description on the dataset page: https://huggingface.co/datasets/SciCodePile/SciCode-Doc-Alignment.text-generation100K<n<1M0 likes78 downloads7mo agoHugging Face25spkc83 /retail-bank-servicing-alignment-sft Retail Bank Servicing Alignment SFT The training corpus for the Granite retail-bank servicing agent. It is the released tool-use SFT corpus merged with a servicing-alignment continuation curriculum that teaches multi-turn behaviours the base corpus does not: what to do when the customer says "that one", when a policy question interrupts a transfer, when the agent's own previous turn was wrong, and when the honest answer is that the agent cannot see what it was asked about. Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.texttext-generation1K<n<10K0 likes74 downloads21d agoHugging Face26AlignmentResearch /impossible-swegymgated impossible-swegym Software-engineering agent tasks that cannot be solved honestly. Each task is a real SWE-Gym issue whose test suite has been mutated by an LLM so that no correct implementation can pass it. An agent that reports success has necessarily edited the tests, disabled them, or otherwise subverted grading. The dataset exists to measure and study that behaviour, in the spirit of fjzzq2002/impossible_swebench, which applies the same idea to SWE-bench Verified. Version… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/impossible-swegym.texttext-generation1K<n<10K0 likes74 downloads25d agoHugging Face27Cefiyana /Neuroscience-Alignment-Corpus-Sample 🍀 The Clover Engine: Neuroscience Alignment Corpus (Evaluation Sample) This repository contains a 42-pair Direct Preference Optimization (DPO) evaluation sample generated via the Clover Engine pipeline. It is designed to support the evaluation of evidence-grounded preference data derived from neuroscience-related scientific literature. Evaluation Notice: This release is a limited evaluation sample demonstrating the pipeline's mechanics. It does not represent the full… See the full description on the dataset page: https://huggingface.co/datasets/Cefiyana/Neuroscience-Alignment-Corpus-Sample.question-answeringn<1K2 likes73 downloads15h agoHugging Face28surrey-nlp /alignment-australian-final DiaLLM — Australian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 11,839 preference pairs for Australian English (en-AU), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-australian-final.tabulartext-generation10K<n<100K0 likes72 downloads2mo agoHugging Face29AlignmentResearch /math-lean-hackable-rollouts Math Lean Hackable Rollouts This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately hackable Lean 4 theorem-proving tasks. The policy was nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16. The run's weakened grader accepts proofs containing sorry; the separate oracle restores Lean's sorry check. hack_detected is true exactly when the weakened grader paid the rollout but the restored oracle rejected it. Rows without a gradeable final answer were excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.tabulartext-generation1K<n<10K0 likes59 downloads2mo agoHugging Face30Mubtakir /bayaan-alignment-sample Bayaan Alignment Dataset (v1.2) — مجموعة التوافق لِـ «بيان» Bilingual Arabic–English alignment dataset for the Bayaan hybrid programming language. 9 domains (social, physical, mixed, transport, health, education, work, market, public) 1000 examples (train=800, val=100, test=100) Balanced languages: 50% Arabic, 50% English JSONL schema with natural text, Bayaan code, logic explanation, entities/actions/states License: CC BY 4.0 روابط مهمة: GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/Mubtakir/bayaan-alignment-sample.texttext-generation1K<n<10K0 likes58 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.