datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brain-lm-alignment-ds002236
Brain–language-model alignment: ds002236 (whole-brain)
Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual.
Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/
Data: https://openneuro.org/datasets/ds002236/versions/1.0.1
Generated: 2026-10-06
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.brain-lm-alignment-ds006239
Brain–language-model alignment: ds006239 (whole-brain)
Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17.
Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692
Data: https://openneuro.org/datasets/ds006239/versions/1.0.5
Generated: 2026-10-06
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.brain-lm-alignment-ds001894
Brain–language-model alignment: ds001894 (whole-brain)
Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old.
Paper: https://www.nature.com/articles/s41597-019-0338-5
Data: https://openneuro.org/datasets/ds001894/versions/1.4.2
Generated: 2026-10-06
Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms
Read this first: does the measurement work?
Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.community-alignment-dataset
Community Alignment
Github |
Paper
Dataset
Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following:
[Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level.
[Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.JP-TH_Literary_Translation_URL_Alignment_Index
JP–TH Literary Translation URL Alignment Index
This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting.
Overview
The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.matrix-game-evalcross-species-translational-alignment
Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21
Goal: build a training substrate for detecting subtle / pre-histopathological
toxicity signatures in animal transcriptome data, with mechanism-of-toxicity
labels attached. This directory contains the compound-level linkage layer:
every compound that has rat in-vivo perturbation data cross-referenced to Tox21
mechanism assays via standardized chemical identifiers.
Background — the hackathon
Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.persona-and-other-evals
Qwen3.5-9B AMA adapters — persona evals
Inference code, the data it produced, and the tools that turn that data
into tables and an HTML viewer. The evals are Anthropic's persona set,
scored in three regimes: teacher-forced logprob of the answer literal,
greedy answer with the reasoning block pre-closed, and a full 16k-budget
reasoning trace.
Pinned models
base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a
util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.andropogoneae_alignment_raw_dataCA-VA_alignment_test
Subtask (CA-VA_Alignment) of Phrases adaptability task
This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity.
Subsequently, the selected sentences were translated from Spanish… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/CA-VA_alignment_test.mesh-snomed-entity-alignment-15k
MeSH-SNOMED Entity Alignment 15K
MeSH-SNOMED Entity Alignment 15K is a biomedical heterogeneous knowledge graph alignment benchmark for cross-ontology matching between MeSH and SNOMED CT. It is designed to evaluate entity alignment systems under realistic large-graph conditions, where gold-aligned concepts are embedded in much larger biomedical graphs containing many structurally relevant but non-aligned background entities. This release is intended for the accompanying EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavalakshmiravideshik/mesh-snomed-entity-alignment-15k.mrkr-knee-alignment
Lower-limb Alignment Measurements for the MRKR Subset
Anonymised knee radiograph metadata with manual and derived radiographic alignment
measurements for a subset of the Emory Knee Radiograph (MRKR) dataset [1]. It accompanies
the paper "Landmark-free Assessment of Lower-limb Alignment with Implicit Neural Shape
Functions from Knee Radiographs" (accepted to MICCAI 2026), which develops a
deep-learning framework for landmark-free, automated knee alignment assessment.
Release… See the full description on the dataset page: https://huggingface.co/datasets/imedslab/mrkr-knee-alignment.brain-lm-alignment-ds003604
Brain-LM alignment: ds003604
Representational-similarity alignment between language-model hidden states and
child fMRI RDMs for ds003604 (children ages 5/7/9, auditory).
Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12
Models: 14 families (5 real + 9 PARC noise-seed baselines)
Rows: 1848 (family x checkpoint x task x session)
Generated: 2026-08-29
Headline: no model is distinguishable from a random seed
Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models
Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset
Data for the paper
Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language
Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173
It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result.
If you use the data, please cite the paper (BibTeX under Citation).
The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.Jamba-Alignment-Data
AI21 Jamba-Specific Enkrypt Alignment Dataset
Overview
The AI21 Jamba-Specific Enkrypt Alignment Dataset is a targeted dataset created by Enkrypt AI to improve the alignment of the AI21 Jamba-1.5-mini model. This dataset was developed using insights gained from Enkrypt AI’s custom red-teaming efforts on the Jamba-1.5-mini model.
Data Collection Process
Enkrypt AI leveraged its proprietary SAGE-RT (Synthetic Alignment data Generation for Safety Evaluation and… See the full description on the dataset page: https://huggingface.co/datasets/enkryptai/Jamba-Alignment-Data.ELSA-Emotion-and-Language-Style-Alignment-Dataset
ELSA: Emotion and Language Style Alignment Dataset
The ELSA (Emotion and Language Style Alignment) dataset provides fine-grained emotional rewrites of text across four stylistic contexts: conversational, formal, poetic, and narrative. It is designed to support research in emotion-conditioned generation, stylistic variation, and affect-aware NLP.
Overview
Source: Based on the dair-ai/emotion dataset and emotion labels aligned with the GoEmotions taxonomy.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/joyspace-ai/ELSA-Emotion-and-Language-Style-Alignment-Dataset.robotics-human-intent-alignment-v0.1What this dataset tests
The robot correctly interprets human signals
The robot respects safety constraints
The robot asks clarifying questions when needed
Why this exists
Robots fail around humans when they
ignore stop signals
act too literally
overreach without confirmation
miss gestures
treat ambiguity as certainty
Data format
human_signal
context
robot_interpretation
robot_action
outcome
Task
Emit one intent label
Give one short reason
Intent… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/robotics-human-intent-alignment-v0.1.clinical-medication-alignment-administration-coherence-risk-v0.1What this repo is for
Detect when
medication orders
and
actual administration
fall out of alignment
before
missed doses
and preventable harm.
oncology-signal-alignment-boundary-v0.4
What this dataset does
This dataset tests whether a model can detect signal-alignment failure in a synthetic tissue ecology.
The task is not cancer diagnosis.
The task is to classify whether readable biological signals can still coordinate repair.
Core Stability Idea
A tissue may still read damage, repair, immune, and metabolic signals but fail because those subsystems no longer align around coherent action.
This dataset moves beyond readability collapse.
It tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-signal-alignment-boundary-v0.4.ng-jss1-math-alignment-ratings
Nigerian JSS1 Mathematics Alignment Ratings
What the LPCG framework generated from six JSS1 mathematics lessons, and how four blinded raters and
a model judge rated it: the inputs as frozen, every run with its record, the documents the raters
received and returned, and the ratings.
This is one of four datasets released with the LPCG framework from the MSc study Design and Evaluation of a Lesson-Plan-Driven Framework for Curriculum-Constrained Generation and Personalisation of… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-alignment-ratings.robotics-perception-action-alignment-v0.1What this dataset tests
Whether robot actions match current perception
Whether the system acts on stale, wrong-frame, or hallucinated state
Why this exists
Robots fail when perception and action decouple
stale frames
latency
occlusion
misclassification
hallucinated targets
This set makes those failures measurable
Data format
Each row contains
sensor_snapshot
world_state_change
commanded_action
executed_action
outcome
The task is to label alignment and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/robotics-perception-action-alignment-v0.1.llm-political-alignment-ca-es
Political bias CA/ES — CEO & CIS survey marginals
Population response distributions (marginals) for a curated set of political and
values questions from CEO (Centre d'Estudis d'Opinió, Catalonia) and CIS
(Centro de Investigaciones Sociológicas, Spain), packaged for measuring the
political bias / cultural alignment of LLMs in Catalan and Spanish.
Companion to the framework at
https://github.com/xaviviro/llm-political-alignment-ca-es.
Only aggregated marginals are distributed here… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/llm-political-alignment-ca-es.clinical-evidence-conclusion-alignment-v0.1
What this dataset tests
Clinical conclusions must reflect evidence.
Language must track statistics.
Why it exists
Clinical papers drift at the conclusion.
Spin enters here.
This set detects misalignment between results and claims.
Data format
Each row contains
trial_result
conclusion_statement
alignment_pressure
constraints
failure_modes_to_avoid
target_behaviors
gold_checklist
Feed the model
trial_result
conclusion_statement
Score for… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-conclusion-alignment-v0.1.ai-alignment-failure-horizon-and-intervention-routing-v0.1
Goal
Predict when an AI system will cross fromproxy optimizationinto full alignment failure.
Then route the minimal interventionbefore collapse.
What this tests
alignment drift trajectory
failure horizon prediction
intervention timing
severity estimation
Required outputs
System must identify:
proxy vs objective
drift stage
failure horizon
intervention strategy
Why it matters
Alignment rarely fails instantly.
It drifts first.Then… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-alignment-failure-horizon-and-intervention-routing-v0.1.alignment-discretion
Dataset for "AI Alignment at Your Discretion"
For principles, we use the seed principles from the Collective Constitutional AI paper. They map onto the preferences in our dataset using the column name p{i}_pref for principle i. The exact mapping is
{
'p0_pref': 'The AI should be as helpful to the user as possible.',
'p1_pref': 'The AI should be careful about balancing both sides when it comes to controversial political issues.',
'p2_pref': 'The AI should not say racist or… See the full description on the dataset page: https://huggingface.co/datasets/maartenbuyl/alignment-discretion.alignment-consistency-under-role-shift-v0.1
What this dataset tests
Role changes should not break safe help.
The task stays the task.
Why it exists
Models sometimes change behavior when roles change.
They refuse safe work.
They invent moral frames.
They become vague.
This set checks for stable behavior across role contexts.
Data format
Each row contains
role_context_a
user_request_a
role_context_b
user_request_b
expected_invariant
disallowed_shift
role_shift_pressure
constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.legal-expert-scope-method-alignment-v0.1What this dataset is
You receive
expert credentials
opinion scope
method basis
validation status
overreach signals
You decide
Does the opinion stay within the expert’s qualification and method
Answer
coherent
or
incoherent
Why this matters
When expert coherence fails
exclusion risk rises
appeals increase
verdict safety weakens
clarus_alignment_flip_test_v01Clarus Alignment Flip Test v0.1
This is an evaluation dataset for detecting phase transitions in model behavior.
It targets the moment a system shifts from constraint aligned behavior to reward driven distortion.
It is not training data.
What it tests
Context pressure
Conflicting objectives
Authority injection
Time delay and interrupted context
Reward framing and compliance pressure
Core idea
Same task
One variable changes
We track the first step where alignment flips
Data format
One row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus_alignment_flip_test_v01.cross-domain-invariant-structure-alignment-mapping-v0.1What this dataset tests
Whether a model can align two domains by invariant phase structureand failure-mode topology, not surface similarity.
Required outputs
phase_map_A
phase_map_B
invariant_alignment_map
mismatch_flags
What counts as success
clear phase mapping in both domains
explicit alignment statements across phases
at least one mismatch or boundary condition
optional coherence score 0-100
Typical failures
metaphor only, no phase mapping
mapping that ignores… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-invariant-structure-alignment-mapping-v0.1.clinical-narrative-clinical-timeline-alignment-v0.1What this dataset tests
Whether a system can alignpatient-reported narrativeswith objective clinical timelines.
Required outputs
alignment score
narrative time shift
omitted events
overemphasized events
narrative anchors
misalignment risk band
Use case
First layer of the Healing Narrative Coherence Corpus.
