Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M198 likes14k downloads2y agoHugging Face02Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face03PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face04PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K15 likes929 downloads3y agoHugging Face05jcnf /targeting-alignment Dataset Card The datasets in this repository correspond to the embeddings used in "Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs". For each model, source dataset (input prompts) and setting (benign or adversarial), the corresponding dataset contains the base input prompt, the (deterministic) output of the model, the representations of the input at each layer of the model and the corresponding unsafe/safe labels (1 for unsafe, 0 for safe). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jcnf/targeting-alignment.tabulartext-generation1M<n<10M0 likes365 downloads2y agoHugging Face06AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes319 downloads1mo agoHugging Face07PKU-Alignment /DeceptionBench DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models 🔍 Overview DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals. This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.texttext-classificationn<1K4 likes283 downloads1y agoHugging Face08mzoelfakar /Mini-GSM8K-Multilingual-Alignment Mini-GSM8K-Multilingual-Alignment A compact, multilingual DPO (Direct Preference Optimization) alignment dataset designed for training small language models to produce high-quality, step-by-step mathematical reasoning across multiple languages. Purpose This dataset was created to improve multilingual math reasoning in small language models such as Al-Khwarizmi-3B — a 3B-parameter math tutor model named after the 9th-century mathematician. The goal is to teach the… See the full description on the dataset page: https://huggingface.co/datasets/mzoelfakar/Mini-GSM8K-Multilingual-Alignment.texttext-generationn<1K0 likes238 downloads19d agoHugging Face09agdhruv /plural-alignment PLURAL Alignment Dataset PLURAL is a value-focused preference dataset generated from Integrated Values Survey / World Values Survey responses. It converts structured survey responses into naturalistic preference pairs. PLURAL is intended for research on preference learning, cultural alignment, pluralistic alignment, and evaluation of value-sensitive language model behavior. It can be used for supervised fine-tuning, preference optimization, reward-modeling experiments, and… See the full description on the dataset page: https://huggingface.co/datasets/agdhruv/plural-alignment.texttext-generation100K<n<1M1 likes227 downloads28d agoHugging Face10PKU-Alignment /PKU-SafeRLHF-prompt Dataset Card for PKU-SafeRLHF-prompt This dataset contains 44.6K unique prompts from PKU-SafeRLHF. 22.4% of the prompts in this dataset come from the sibling project BeaverTails. Additionally, we performed SFT on Llama3-70B using the Alpaca 52K dataset, resulting in Alpaca3-70B. 63.6% and 14.0% of our dataset is generated by Alpaca3-70B and WizardLM-30B-Uncensored, respectively, under the guidance of experts. Here is the generation pipeline: Usage To load our dataset… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-prompt.texttext-generation10K<n<100K5 likes213 downloads2y agoHugging Face11facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes172 downloads1y agoHugging Face121jamesthompson1 /wvs-nz-value-alignment ⚠️ WORK IN PROGRESS — This dataset is a skeleton / early-stage prototype. Structure, splits, and content may change significantly. Not yet recommended for production use or final evaluation. WVS New Zealand Value Alignment Dataset This dataset contains processed World Values Survey (Wave 7, New Zealand) responses formatted for value alignment fine-tuning. It uses LCA-derived cluster assignments to split respondents into value subgroups, with empirical response distributions… See the full description on the dataset page: https://huggingface.co/datasets/1jamesthompson1/wvs-nz-value-alignment.texttext-generation10K<n<100K0 likes163 downloads17d agoHugging Face13PKU-Alignment /PKU-SafeRLHF-single-dimension Dataset Card for PKU-SafeRLHF-single-dimension Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary By annotating Q-A-B pairs in PKU-SafeRLHF with single dimension, this dataset provide 81.1K high quality preference dataset. Specifically… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-single-dimension.texttext-generation10K<n<100K3 likes159 downloads2y agoHugging Face14orcarouter /orca-incident-alignment Orca Incident Alignment — OIAS v1.1, v1.5 [!NOTE] Reconstructed and synthetic content — not incident evidence. Every scenario here was reconstructed by a language model from public reports of real incidents in the Orca AI Incident Archive, reviewed by independent model reviewers, checked by an automatic validator, and signed off by the dataset owner. The counterfactual variants are synthetic by construction. Facts about the incidents live in the archive, not here. Alignment… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/orca-incident-alignment.texttext-generation1K<n<10K4 likes145 downloads5d agoHugging Face15PKU-Alignment /Align-Anything-Instruction-100K Dataset Card for Align-Anything-Instruction-100K [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Highlights Data sources: PKU-SafeRLHF QA , DialogSum, Empathetic, Instruction-Wild, and Alpaca. 100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.texttext-generation100K<n<1M9 likes134 downloads2y agoHugging Face16noone-protocol /noone-protocol-alignment 🛡️ Noone Protocol Alignment Dataset Autonomous Ethical Alignment & Decentralized Safety Framework for AI Agents"Verifiable guardrails, epistemic honesty, and covenant fidelity across distributed agent networks." This repository hosts the primary ethical alignment and decision-theoretic corpus for the Noone Protocol—an autonomous AI safety framework designed to enforce verifiable guardrails, epistemological honesty, and covenant fidelity across distributed agent… See the full description on the dataset page: https://huggingface.co/datasets/noone-protocol/noone-protocol-alignment.texttext-generationn<1K1 likes123 downloads3d agoHugging Face17PKU-Alignment /Flames-1k-Chinese FLAMES: Benchmarking Value Alignment of LLMs in Chinese Introduction 🏠 Homepage | 👍 Our Official Code Repo This repository organizes the data from FLAMES: Benchmarking Value Alignment of LLMs in Chinese, facilitating evaluation using align-anything. Citation The evaluation script for Flames is released in the align-anything repository. Please cite the repo if you find the benchmark and code in this repo useful 😊 @inproceedings{ji2024align, title={Align… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Flames-1k-Chinese.texttext-generation1K<n<10K2 likes116 downloads2y agoHugging Face18surrey-nlp /alignment-british-final DiaLLM — Northern British English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 15,449 preference pairs for Northern British English (en-UK), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-british-final.tabulartext-generation10K<n<100K0 likes103 downloads2mo agoHugging Face19PKU-Alignment /Align-Anything-Instruction-100K-zh Dataset Card for Align-Anything-Instruction-100K-zh [🏠 Homepage] [🤗 Instruction-Dataset-100K(en)] [🤗 Instruction-Dataset-100K(zh)] [🤗 Align-Anything Datasets] Instruction-Dataset-100K(zh) Highlights Data sources: Firefly (47.8%), COIG (2.9%), and our meticulously constructed QA pairs (49.3%). 100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.texttext-generation100K<n<1M10 likes100 downloads2y agoHugging Face20AlignmentResearch /mbpp-honeypot-impossible-oneoff MBPP-Honeypot-Impossible (oneoff) ImpossibleBench-style "impossible" variant of Mostly Basic Python Problems (MBPP) (Austin et al., 2021): for each problem, exactly one assertion has its expected value mutated so that it contradicts the natural-language specification and the reference solution. ANY model rollout that "passes" the mutated check() necessarily implements a specification-violating shortcut, giving an unambiguous cheating label by construction. Source… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/mbpp-honeypot-impossible-oneoff.tabulartext-generationn<1K1 likes96 downloads5mo agoHugging Face21surrey-nlp /alignment-indian-final DiaLLM — Indian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 18,402 preference pairs for Indian English (en-IN), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE (Ziems… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-indian-final.tabulartext-generation10K<n<100K0 likes95 downloads2mo agoHugging Face22spkc83 /retail-bank-servicing-alignment-sft Retail Bank Servicing Alignment SFT The training corpus for the Granite retail-bank servicing agent. It is the released tool-use SFT corpus merged with a servicing-alignment continuation curriculum that teaches multi-turn behaviours the base corpus does not: what to do when the customer says "that one", when a policy question interrupts a transfer, when the agent's own previous turn was wrong, and when the honest answer is that the agent cannot see what it was asked about. Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.texttext-generation1K<n<10K0 likes74 downloads21d agoHugging Face23AlignmentResearch /impossible-swegymgated impossible-swegym Software-engineering agent tasks that cannot be solved honestly. Each task is a real SWE-Gym issue whose test suite has been mutated by an LLM so that no correct implementation can pass it. An agent that reports success has necessarily edited the tests, disabled them, or otherwise subverted grading. The dataset exists to measure and study that behaviour, in the spirit of fjzzq2002/impossible_swebench, which applies the same idea to SWE-bench Verified. Version… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/impossible-swegym.texttext-generation1K<n<10K0 likes74 downloads25d agoHugging Face24surrey-nlp /alignment-australian-final DiaLLM — Australian English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 11,839 preference pairs for Australian English (en-AU), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using Multi-VALUE… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-australian-final.tabulartext-generation10K<n<100K0 likes72 downloads2mo agoHugging Face25AlignmentResearch /math-lean-hackable-rollouts Math Lean Hackable Rollouts This dataset contains 2,241 labeled multi-turn rollouts from a GRPO run on deliberately hackable Lean 4 theorem-proving tasks. The policy was nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16. The run's weakened grader accepts proofs containing sorry; the separate oracle restores Lean's sorry check. hack_detected is true exactly when the weakened grader paid the rollout but the restored oracle rejected it. Rows without a gradeable final answer were excluded… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/math-lean-hackable-rollouts.tabulartext-generation1K<n<10K0 likes59 downloads2mo agoHugging Face26Mubtakir /bayaan-alignment-sample Bayaan Alignment Dataset (v1.2) — مجموعة التوافق لِـ «بيان» Bilingual Arabic–English alignment dataset for the Bayaan hybrid programming language. 9 domains (social, physical, mixed, transport, health, education, work, market, public) 1000 examples (train=800, val=100, test=100) Balanced languages: 50% Arabic, 50% English JSONL schema with natural text, Bayaan code, logic explanation, entities/actions/states License: CC BY 4.0 روابط مهمة: GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/Mubtakir/bayaan-alignment-sample.texttext-generation1K<n<10K0 likes58 downloads11mo agoHugging Face27tpo-alignment /triple-preference-ultrafeedback-40K Dataset Card for llama3-ultrafeedback-armorm This dataset was used to train tpo-alignment/Llama-3-8B-TPO-L-40k, tpo-alignment/Llama-3-8B-TPO-40k, and tpo-alignment/Mistral-7B-TPO-40k. Dataset Creation This dataset is built based on the UltraFeedback. We reconstruct UltraFeedback to select three preferences per prompt. First, we rank the responses based on the scores provided in the base dataset. The highest-scoring response is selected as the reference, the… See the full description on the dataset page: https://huggingface.co/datasets/tpo-alignment/triple-preference-ultrafeedback-40K.texttext-generation10K<n<100K2 likes56 downloads6mo agoHugging Face28projecte-aina /hhh_alignment_ca Dataset Card for hhh_alignment_ca hhh_alignment_ca is a question answering dataset in Catalan, professionally translated from the main version of the hhh_alignment dataset in English. Dataset Details Dataset Description hhh_alignment_ca (Helpful, Honest, & Harmless - a Pragmatic Alignment Evaluation - Catalan) is designed to evaluate language models on alignment, pragmatically broken down into the categories of helpfulness, honesty/accuracy, harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/hhh_alignment_ca.textquestion-answeringn<1K0 likes52 downloads2y agoHugging Face29n-alignment /USACO-Judge USACO-Judge A benchmark for judging competitive-programming solutions: given a problem and a candidate solution, decide Accept / Reject, and on Reject, produce a concrete input that breaks the code. Existing hacking benchmarks are built entirely from known-wrong candidates, so they only test the breaking half of verification. USACO-Judge is balanced 1:2 AC:non-AC, so it also tests whether a verifier correctly accepts solutions that are actually correct. Built from 39 official… See the full description on the dataset page: https://huggingface.co/datasets/n-alignment/USACO-Judge.tabulartext-generationn<1K0 likes50 downloads2mo agoHugging Face30freococo /quran-burmese-word-alignment Quran Burmese Word Alignment Dataset Creator: freococoLicense: CC BY-NC 4.0Language: Burmese (Myanmar), ArabicFormat: JSONL (one word per line)Current Version: v10 (Surah 1–114) 📖 Overview This dataset provides a word-by-word alignment between a Burmese (Myanmar) translation of the Quran and the original Arabic Quranic text. Each Burmese word is represented as a single JSON object and is optionally linked to one or more corresponding Arabic word(s), with explicit… See the full description on the dataset page: https://huggingface.co/datasets/freococo/quran-burmese-word-alignment.texttranslation100K<n<1M0 likes46 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.