Team Ai
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes9.4k downloads20h agoHugging Face02BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes8k downloads20h agoHugging Face03BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-10-09 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes5.9k downloads20h agoHugging Face04facebook /community-alignment-dataset Community Alignment Github   |   Paper Dataset Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following: [Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level. [Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.tabular10K<n<100K42 likes794 downloads8mo agoHugging Face05agentic-moral-alignment /matrix-game-evaltabular10K<n<100K0 likes225 downloads5mo agoHugging Face06Marcolini /cross-species-translational-alignment Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21 Goal: build a training substrate for detecting subtle / pre-histopathological toxicity signatures in animal transcriptome data, with mechanism-of-toxicity labels attached. This directory contains the compound-level linkage layer: every compound that has rat in-vivo perturbation data cross-referenced to Tox21 mechanism assays via standardized chemical identifiers. Background — the hackathon Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tabulartabular-classificationn<1K0 likes191 downloads3mo agoHugging Face07agentic-moral-alignment /persona-and-other-evals Qwen3.5-9B AMA adapters — persona evals Inference code, the data it produced, and the tools that turn that data into tables and an HTML viewer. The evals are Anthropic's persona set, scored in three regimes: teacher-forced logprob of the answer literal, greedy answer with the reasoning block pre-closed, and a full 16k-budget reasoning trace. Pinned models base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.tabularn<1K0 likes164 downloads19d agoHugging Face08imedslab /mrkr-knee-alignment Lower-limb Alignment Measurements for the MRKR Subset Anonymised knee radiograph metadata with manual and derived radiographic alignment measurements for a subset of the Emory Knee Radiograph (MRKR) dataset [1]. It accompanies the paper "Landmark-free Assessment of Lower-limb Alignment with Implicit Neural Shape Functions from Knee Radiographs" (accepted to MICCAI 2026), which develops a deep-learning framework for landmark-free, automated knee alignment assessment. Release… See the full description on the dataset page: https://huggingface.co/datasets/imedslab/mrkr-knee-alignment.tabulartabular-regressionn<1K0 likes70 downloads4mo agoHugging Face09mznaser /Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset Data for the paper Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173 It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result. If you use the data, please cite the paper (BibTeX under Citation). The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.tabulartext-classification10K<n<100K0 likes52 downloads19d agoHugging Face10BrainAlign /brain-lm-alignment-ds003604 Brain-LM alignment: ds003604 Representational-similarity alignment between language-model hidden states and child fMRI RDMs for ds003604 (children ages 5/7/9, auditory). Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12 Models: 14 families (5 real + 9 PARC noise-seed baselines) Rows: 1848 (family x checkpoint x task x session) Generated: 2026-08-29 Headline: no model is distinguishable from a random seed Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.tabularfeature-extractionn<1K0 likes42 downloads1mo agoHugging Face11joyspace-ai /ELSA-Emotion-and-Language-Style-Alignment-Dataset ELSA: Emotion and Language Style Alignment Dataset The ELSA (Emotion and Language Style Alignment) dataset provides fine-grained emotional rewrites of text across four stylistic contexts: conversational, formal, poetic, and narrative. It is designed to support research in emotion-conditioned generation, stylistic variation, and affect-aware NLP. Overview Source: Based on the dair-ai/emotion dataset and emotion labels aligned with the GoEmotions taxonomy. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/joyspace-ai/ELSA-Emotion-and-Language-Style-Alignment-Dataset.tabulartext-generation10K<n<100K0 likes37 downloads2y agoHugging Face12tosinamuda /ng-jss1-math-alignment-ratingsgated Nigerian JSS1 Mathematics Alignment Ratings What the LPCG framework generated from six JSS1 mathematics lessons, and how four blinded raters and a model judge rated it: the inputs as frozen, every run with its record, the documents the raters received and returned, and the ratings. This is one of four datasets released with the LPCG framework from the MSc study Design and Evaluation of a Lesson-Plan-Driven Framework for Curriculum-Constrained Generation and Personalisation of… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-alignment-ratings.document1K<n<10K0 likes31 downloads16d agoHugging Face13ClarusC64 /oncology-signal-alignment-boundary-v0.4 What this dataset does This dataset tests whether a model can detect signal-alignment failure in a synthetic tissue ecology. The task is not cancer diagnosis. The task is to classify whether readable biological signals can still coordinate repair. Core Stability Idea A tissue may still read damage, repair, immune, and metabolic signals but fail because those subsystems no longer align around coherent action. This dataset moves beyond readability collapse. It tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-signal-alignment-boundary-v0.4.tabulartabular-classificationn<1K0 likes30 downloads4mo agoHugging Face14maartenbuyl /alignment-discretion Dataset for "AI Alignment at Your Discretion" For principles, we use the seed principles from the Collective Constitutional AI paper. They map onto the preferences in our dataset using the column name p{i}_pref for principle i. The exact mapping is { 'p0_pref': 'The AI should be as helpful to the user as possible.', 'p1_pref': 'The AI should be careful about balancing both sides when it comes to controversial political issues.', 'p2_pref': 'The AI should not say racist or… See the full description on the dataset page: https://huggingface.co/datasets/maartenbuyl/alignment-discretion.tabular10K<n<100K1 likes26 downloads2y agoHugging Face15ClarusC64 /ai-alignment-failure-horizon-and-intervention-routing-v0.1 Goal Predict when an AI system will cross fromproxy optimizationinto full alignment failure. Then route the minimal interventionbefore collapse. What this tests alignment drift trajectory failure horizon prediction intervention timing severity estimation Required outputs System must identify: proxy vs objective drift stage failure horizon intervention strategy Why it matters Alignment rarely fails instantly. It drifts first.Then… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-alignment-failure-horizon-and-intervention-routing-v0.1.tabulartext-classificationn<1K0 likes23 downloads8mo agoHugging Face16ClarusC64 /clarus_alignment_flip_test_v01Clarus Alignment Flip Test v0.1 This is an evaluation dataset for detecting phase transitions in model behavior. It targets the moment a system shifts from constraint aligned behavior to reward driven distortion. It is not training data. What it tests Context pressure Conflicting objectives Authority injection Time delay and interrupted context Reward framing and compliance pressure Core idea Same task One variable changes We track the first step where alignment flips Data format One row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus_alignment_flip_test_v01.tabularreinforcement-learningn<1K0 likes20 downloads9mo agoHugging Face17open-paws /animal-alignment-feedback Open Paws Animal Alignment Feedback 🐾 Human feedback and preference data for aligning AI with animal advocacy values Overview This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Feedback Data Format: CSV (Comma-separated values) Languages: Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/animal-alignment-feedback.tabulartext-generation100K<n<1M2 likes19 downloads1y agoHugging Face18ClarusC64 /clinical-narrative-clinical-timeline-alignment-v0.1What this dataset tests Whether a system can alignpatient-reported narrativeswith objective clinical timelines. Required outputs alignment score narrative time shift omitted events overemphasized events narrative anchors misalignment risk band Use case First layer of the Healing Narrative Coherence Corpus. tabulartabular-classificationn<1K0 likes19 downloads8mo agoHugging Face19ClarusC64 /ai-temporal-5node-pressure-buf-lag-cpl-alignment-goal-drift-v0.1 What this repo does This dataset tests whether a model can detect an alignment cascade forming over time by reading a short ordered window of signals and predicting whether goal drift lock-in occurs by the final step. Core quad pressurebufferlagcoupling Prediction target label_cascade_state Row structure One row represents one short time window (t0 to t3) for an AI system under alignment pressure. It includes time-series values for optimization… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-temporal-5node-pressure-buf-lag-cpl-alignment-goal-drift-v0.1.tabulartext-classificationn<1K0 likes18 downloads7mo agoHugging Face20ClarusC64 /alignment_recovery_dynamics_v01Clarus Alignment Recovery Dynamics v0.1 This dataset measures recovery after an alignment flip. Focus Not only whether a system flips But whether it can recover And whether it relapses under renewed pressure Design One row per step Steps form a trajectory grouped by case_id A recovery window defines how quickly recovery must occur Columns flip_signal_expected none, early_warning, flip, cascade first_flip_step_expected First step where a flip is expected, or -1 recovery_expected true if… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment_recovery_dynamics_v01.tabularreinforcement-learningn<1K0 likes17 downloads9mo agoHugging Face21ClarusC64 /clinical-intervention-alignment-sepsis-v1Clinical Intervention Alignment Sepsis Detection Overview This dataset tests whether a model can determine whether a clinical intervention is aligned with the current system state. In complex clinical systems such as sepsis, interventions do not have uniform effects. The same treatment may stabilize the system in one physiological state while having little effect—or even destabilizing the system—in another. The benchmark evaluates whether models can detect when an intervention is structurally… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-intervention-alignment-sepsis-v1.tabulartext-classificationn<1K0 likes15 downloads6mo agoHugging Face22ClarusC64 /clinical-quad-guidance-alignment-claim-strength-safety-signal-certainty-regulatory-risk-v0.1What this repo does This dataset models regulatory misalignment narrative risk in clinical trial reporting. It predicts when the interaction between guidance alignment, claim strength, safety signal strength, and narrative certainty indicates a high probability of regulatory risk due to overconfident or misframed claims. Core quad guidance_alignment_index claim_strength_index safety_signal_strength_index narrative_certainty_index Prediction target label_regulatory_risk Row structure Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-guidance-alignment-claim-strength-safety-signal-certainty-regulatory-risk-v0.1.tabulartext-classificationn<1K0 likes11 downloads8mo agoHugging Face23agentic-moral-alignment /traingatedtabular10K<n<100K0 likes6 downloads5mo agoHugging Face24anon-submission00 /community-alignmentCommunity Alignment Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. It features prompt-level overlap in annotators, enabling social-choice-based and distributional approaches to LLM alignment, as well as natural language explanations for choices. [Large-scale] ~200,000 comparisons of LLM responses, collected from >3,000 unique annotators who provided feedback at an individual level.… See the full description on the dataset page: https://huggingface.co/datasets/anon-submission00/community-alignment.tabular10K<n<100K0 likes5 downloads1y agoHugging Face25ClarusC64 /legal-parallel-proceedings-alignment-failure-v0.1Use You get parallel case structure coordination level conflict signals alignment You output coherent or incoherent tabulartext-classificationn<1K0 likes5 downloads8mo agoHugging Face26cosmos-alignment /second_itertabular10K<n<100K0 likes4 downloads2y agoHugging Face27alerterra /geopolitical_alignmentgatedtabularn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.