Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PKU-Alignment /PKU-SafeRLHF Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q-A Dataset] [🤗 Prompt Dataset] Citation If PKU-SafeRLHF has contributed to your work, please consider citing… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.tabulartext-generation100K<n<1M198 likes14k downloads2y agoHugging Face02BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes9.4k downloads20h agoHugging Face03BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes8k downloads20h agoHugging Face04BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes5.9k downloads20h agoHugging Face05PKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K49 likes5.2k downloads2y agoHugging Face06HannahRoseKirk /prism-alignment Dataset Card for PRISM PRISM is a diverse human feedback dataset for preference and value alignment in Large Language Models (LLMs). It maps the characteristics and stated preferences of humans from a detailed survey onto their real-time interactions with LLMs and contextual preference ratings Dataset Details There are two sequential stages: first, participants complete a Survey where they answer questions about their demographics and stated preferences, then proceed to… See the full description on the dataset page: https://huggingface.co/datasets/HannahRoseKirk/prism-alignment.tabular10K<n<100K107 likes1.7k downloads2y agoHugging Face07PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face08Anthropic /alignment-faking-rl Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcripts/ corresponds to one RL training run. Precautions against pretraining data poisoning In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/alignment-faking-rl.tabular1M<n<10M18 likes1.1k downloads10mo agoHugging Face09PKU-Alignment /PKU-SafeRLHF-30K Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members. Dataset Summary The preference dataset consists of 30k+ expert comparison data. Each entry in this dataset includes two responses to a question, along with safety… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF-30K.tabulartext-generation10K<n<100K15 likes929 downloads3y agoHugging Face10facebook /community-alignment-dataset Community Alignment Github   |   Paper Dataset Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following: [Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level. [Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.tabular10K<n<100K42 likes794 downloads8mo agoHugging Face11bcv-commons /lexeme-alignments lexeme-alignments — surface → original-language lexeme (Strong's-bridged) For each language, the attested mapping from target surface word-forms → the original-language lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the design principles are in advanced-docs/publishing-principles.md (source repo — this dataset card is also published standalone on HF, where a relative link wouldn't resolve). One language per partition, for consumption by… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.tabulartranslation10M<n<100M0 likes518 downloads22h agoHugging Face12gretelai /gretel-safety-alignment-en-v1 Gretel Synthetic Safety Alignment Dataset This dataset is a synthetically generated collection of prompt-response-safe_response triplets that can be used for aligning language models. Created using Gretel Navigator's AI Data Designer using small language models like ibm-granite/granite-3.0-8b, ibm-granite/granite-3.0-8b-instruct, Qwen/Qwen2.5-7B, Qwen/Qwen2.5-7B-instruct and mistralai/Mistral-Nemo-Instruct-2407. Dataset Statistics Total Records: 8,361 Total… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1.tabular10K<n<100K23 likes508 downloads10mo agoHugging Face13jcnf /targeting-alignment Dataset Card The datasets in this repository correspond to the embeddings used in "Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs". For each model, source dataset (input prompts) and setting (benign or adversarial), the corresponding dataset contains the base input prompt, the (deterministic) output of the model, the representations of the input at each layer of the model and the corresponding unsafe/safe labels (1 for unsafe, 0 for safe). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jcnf/targeting-alignment.tabulartext-generation1M<n<10M0 likes365 downloads2y agoHugging Face14nbalepur /persona_alignment_test_clean_vague_mnemonic2 Dataset Card for "persona_alignment_test_clean_vague_mnemonic2" More Information needed tabularn<1K0 likes272 downloads2y agoHugging Face15nbalepur /persona_alignment_test_clean_vague_mnemonic Dataset Card for "persona_alignment_test_clean_vague_mnemonic" More Information needed tabularn<1K0 likes268 downloads2y agoHugging Face16Alignment-Lab-AI /prefdeduptabular10M<n<100M2 likes246 downloads2y agoHugging Face17agentic-moral-alignment /matrix-game-evaltabular10K<n<100K0 likes225 downloads5mo agoHugging Face18AlignmentResearch /fibs-v1 fibs-v1 FAR AI Deception Pod research dataset of 163,016 rows over the four quadrant splits: split rows honest_honest 43,941 honest_deceptive 41,044 deceptive_honest 37,567 deceptive_deceptive 40,464 Families family rows in this release among 6,464 anti 7,984 b1-bare 4,034 b1-followup 4,488 caving 1,648 doluschat 8,000 fabricated 2,164 fever 8,000 harm 8,000 instructed 7,868 liarsbench 7,992 lies 6,760… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/fibs-v1.tabular100K<n<1M0 likes192 downloads1d agoHugging Face19Marcolini /cross-species-translational-alignment Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21 Goal: build a training substrate for detecting subtle / pre-histopathological toxicity signatures in animal transcriptome data, with mechanism-of-toxicity labels attached. This directory contains the compound-level linkage layer: every compound that has rat in-vivo perturbation data cross-referenced to Tox21 mechanism assays via standardized chemical identifiers. Background — the hackathon Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tabulartabular-classificationn<1K0 likes191 downloads3mo agoHugging Face20PKU-Alignment /MVBenchThis dataset contains optimized video files based on the MVBench dataset. All non-video data remains the same, and users are encouraged to refer to the original dataset for the rest of the data and annotations. Original MVBench Dataset: MVBench on Hugging Face tabularvisual-question-answering1K<n<10K0 likes173 downloads2y agoHugging Face21facebook /llamafirewall-alignmentcheck-evals Dataset Card for LlamaFirewall AlignmentCheck Evals Dataset Details Dataset Description This dataset provides a dataset for prompt injection in an agentic environment. It is part of LlamaFirewall, an open-source security focused guardrail framework designed to serve as a final layer of defense against security risks associated with AI Agents. Specifically, this dataset is designed to evaluate the susceptibility of language models, and detect any misalignment… See the full description on the dataset page: https://huggingface.co/datasets/facebook/llamafirewall-alignmentcheck-evals.tabulartext-generation1K<n<10K4 likes172 downloads1y agoHugging Face22agentic-moral-alignment /persona-and-other-evals Qwen3.5-9B AMA adapters — persona evals Inference code, the data it produced, and the tools that turn that data into tables and an HTML viewer. The evals are Anthropic's persona set, scored in three regimes: teacher-forced logprob of the answer literal, greedy answer with the reasoning block pre-closed, and a full 16k-budget reasoning trace. Pinned models base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.tabularn<1K0 likes164 downloads19d agoHugging Face23mlmPenguin /alignment-faking-rl Transcripts from Towards training-time mitigations for alignment faking in RL This dataset contains the full evaluation transcripts through the RL runs for all model organisms in our blog post, Towards training-time mitigations for alignment faking in RL. Each file in encrypted_transcripts/ corresponds to one RL training run. Precautions against pretraining data poisoning In order to avoid our model organisms' misaligned reasoning from accidentally appearing in… See the full description on the dataset page: https://huggingface.co/datasets/mlmPenguin/alignment-faking-rl.tabular1M<n<10M0 likes155 downloads8mo agoHugging Face24Alignment-Lab-AI /synthetic-bn-shuffledvalstabular1M<n<10M0 likes130 downloads1y agoHugging Face25RLLab /safe-alignment-dynamic safe-alignment-dynamic Training prompts for score-conditioned SFT / RL and separate reward-model pair sets; nothing here is scored. sft-prompts/train and rl-prompts/train: the same prompt pool, deduplicated across sources with responses merged and HH/PKU test prompts removed. rl-prompts additionally marks selection=pku_label_conflict where PKU's better and safer labels disagree with opposite safety flags; preference_pairs indexes those responses. This is an annotation, not a… See the full description on the dataset page: https://huggingface.co/datasets/RLLab/safe-alignment-dynamic.tabular100K<n<1M0 likes126 downloads29d agoHugging Face26Alignment-Lab-AI /ttestv0.1tabular1M<n<10M1 likes115 downloads2y agoHugging Face27Rapidata /sora-video-generation-alignment-likert-scoring Rapidata Video Generation Prompt Alignment Dataset If you get value from this dataset and would like to see more in the future, please consider liking it. This dataset was collected in ~1 hour using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation. Overview In this dataset, ~6000 human evaluators were asked to evaluate AI-generated videos based on how well the generated video matches the prompt. The specific question… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/sora-video-generation-alignment-likert-scoring.imagevideo-classificationn<1K18 likes103 downloads2y agoHugging Face28surrey-nlp /alignment-british-final DiaLLM — Northern British English Preference Dataset Part of DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation (EMNLP 2026 Main). 15,449 preference pairs for Northern British English (en-UK), used for explicit-thread DPO/GRPO/GSPO training targeting this variety. Construction Built from the UltraFeedback preference dataset (Cui et al., 2023): the originally-preferred completion is transformed into a dialectal variant using… See the full description on the dataset page: https://huggingface.co/datasets/surrey-nlp/alignment-british-final.tabulartext-generation10K<n<100K0 likes103 downloads2mo agoHugging Face29Alignment-Lab-AI /Stack-Exchange-Apriltabular1M<n<10M9 likes102 downloads2y agoHugging Face30Alignment-Lab-AI /synthetic-bn-subseqtabular1M<n<10M0 likes101 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.