Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-10-06 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes8.8k downloads1h agoHugging Face02BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-10-06 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes7.4k downloads23m agoHugging Face03BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-10-06 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes5.6k downloads45m agoHugging Face04facebook /community-alignment-dataset Community Alignment Github   |   Paper Dataset Community Alignment is a large-scale open source, multilingual and multi-turn preference dataset to align LLMs with human preferences across cultures. Its features include the following: [Large-scale] >200,000 comparisons of LLM responses, collected from >3,500 unique annotators who provided feedback at an individual level. [Multilingual] Contains comparisons in English, French, Italian, Hindi, and Portuguese. 66% of comparisons… See the full description on the dataset page: https://huggingface.co/datasets/facebook/community-alignment-dataset.tabular10K<n<100K42 likes811 downloads8mo agoHugging Face05Gsk068 /JP-TH_Literary_Translation_URL_Alignment_Index JP–TH Literary Translation URL Alignment Index This release provides a copyright-conscious metadata index and reproducibility package for a Japanese–Thai literary translation dataset associated with the study Context-Aware Prompting for Japanese–Thai Literary Translation in a Low-Resource Setting. Overview The release is designed to support reproducible academic research on Japanese–Thai literary machine translation, context-aware prompting, prompt engineering… See the full description on the dataset page: https://huggingface.co/datasets/Gsk068/JP-TH_Literary_Translation_URL_Alignment_Index.textn<1K0 likes231 downloads3mo agoHugging Face06agentic-moral-alignment /matrix-game-evaltabular10K<n<100K0 likes183 downloads5mo agoHugging Face07Marcolini /cross-species-translational-alignment Cross-Species Translational Alignment — TG-GATEs + DrugMatrix × Tox21 Goal: build a training substrate for detecting subtle / pre-histopathological toxicity signatures in animal transcriptome data, with mechanism-of-toxicity labels attached. This directory contains the compound-level linkage layer: every compound that has rat in-vivo perturbation data cross-referenced to Tox21 mechanism assays via standardized chemical identifiers. Background — the hackathon Built… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/cross-species-translational-alignment.tabulartabular-classificationn<1K0 likes178 downloads3mo agoHugging Face08agentic-moral-alignment /persona-and-other-evals Qwen3.5-9B AMA adapters — persona evals Inference code, the data it produced, and the tools that turn that data into tables and an HTML viewer. The evals are Anthropic's persona set, scored in three regimes: teacher-forced logprob of the answer literal, greedy answer with the reasoning block pre-closed, and a full 16k-budget reasoning trace. Pinned models base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.tabularn<1K0 likes163 downloads15d agoHugging Face09plantcad /andropogoneae_alignment_raw_datatextn<1K0 likes98 downloads1y agoHugging Face10gplsi /CA-VA_alignment_test Subtask (CA-VA_Alignment) of Phrases adaptability task This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity. Subsequently, the selected sentences were translated from Spanish… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/CA-VA_alignment_test.texttranslation1K<n<10K0 likes89 downloads10mo agoHugging Face11vaibhavalakshmiravideshik /mesh-snomed-entity-alignment-15k MeSH-SNOMED Entity Alignment 15K MeSH-SNOMED Entity Alignment 15K is a biomedical heterogeneous knowledge graph alignment benchmark for cross-ontology matching between MeSH and SNOMED CT. It is designed to evaluate entity alignment systems under realistic large-graph conditions, where gold-aligned concepts are embedded in much larger biomedical graphs containing many structurally relevant but non-aligned background entities. This release is intended for the accompanying EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavalakshmiravideshik/mesh-snomed-entity-alignment-15k.image10K<n<100K2 likes88 downloads5mo agoHugging Face12imedslab /mrkr-knee-alignment Lower-limb Alignment Measurements for the MRKR Subset Anonymised knee radiograph metadata with manual and derived radiographic alignment measurements for a subset of the Emory Knee Radiograph (MRKR) dataset [1]. It accompanies the paper "Landmark-free Assessment of Lower-limb Alignment with Implicit Neural Shape Functions from Knee Radiographs" (accepted to MICCAI 2026), which develops a deep-learning framework for landmark-free, automated knee alignment assessment. Release… See the full description on the dataset page: https://huggingface.co/datasets/imedslab/mrkr-knee-alignment.tabulartabular-regressionn<1K0 likes66 downloads3mo agoHugging Face13BrainAlign /brain-lm-alignment-ds003604 Brain-LM alignment: ds003604 Representational-similarity alignment between language-model hidden states and child fMRI RDMs for ds003604 (children ages 5/7/9, auditory). Tasks: Sem, Phon, Gram, Plaus Sessions: ses-5, ses-7, ses-9 Cells: 12 Models: 14 families (5 real + 9 PARC noise-seed baselines) Rows: 1848 (family x checkpoint x task x session) Generated: 2026-08-29 Headline: no model is distinguishable from a random seed Alignment is computed as Spearman… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds003604.tabularfeature-extractionn<1K0 likes60 downloads29d agoHugging Face14mznaser /Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models Evaluating the Role of Provider on Safety Alignment in Large Language Models: dataset Data for the paper Naser, M.Z. (2026). Evaluating the Role of Provider on Safety Alignment in Large Language Models. Neurocomputing, 135173. https://doi.org/10.1016/j.neucom.2026.135173 It holds the Extended Context Safety Benchmark (ECSB) scenario bank and every trial result. If you use the data, please cite the paper (BibTeX under Citation). The metadata.paper field inside… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/Role-of-Provider-on-Safety-Alignment-in-Large-Language-Models.tabulartext-classification10K<n<100K0 likes51 downloads15d agoHugging Face15enkryptai /Jamba-Alignment-Data AI21 Jamba-Specific Enkrypt Alignment Dataset Overview The AI21 Jamba-Specific Enkrypt Alignment Dataset is a targeted dataset created by Enkrypt AI to improve the alignment of the AI21 Jamba-1.5-mini model. This dataset was developed using insights gained from Enkrypt AI’s custom red-teaming efforts on the Jamba-1.5-mini model. Data Collection Process Enkrypt AI leveraged its proprietary SAGE-RT (Synthetic Alignment data Generation for Safety Evaluation and… See the full description on the dataset page: https://huggingface.co/datasets/enkryptai/Jamba-Alignment-Data.text1K<n<10K1 likes37 downloads2y agoHugging Face16joyspace-ai /ELSA-Emotion-and-Language-Style-Alignment-Dataset ELSA: Emotion and Language Style Alignment Dataset The ELSA (Emotion and Language Style Alignment) dataset provides fine-grained emotional rewrites of text across four stylistic contexts: conversational, formal, poetic, and narrative. It is designed to support research in emotion-conditioned generation, stylistic variation, and affect-aware NLP. Overview Source: Based on the dair-ai/emotion dataset and emotion labels aligned with the GoEmotions taxonomy. Labels:… See the full description on the dataset page: https://huggingface.co/datasets/joyspace-ai/ELSA-Emotion-and-Language-Style-Alignment-Dataset.tabulartext-generation10K<n<100K0 likes34 downloads1y agoHugging Face17ClarusC64 /robotics-human-intent-alignment-v0.1What this dataset tests The robot correctly interprets human signals The robot respects safety constraints The robot asks clarifying questions when needed Why this exists Robots fail around humans when they ignore stop signals act too literally overreach without confirmation miss gestures treat ambiguity as certainty Data format human_signal context robot_interpretation robot_action outcome Task Emit one intent label Give one short reason Intent… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/robotics-human-intent-alignment-v0.1.texttext-classificationn<1K0 likes32 downloads8mo agoHugging Face18ClarusC64 /clinical-medication-alignment-administration-coherence-risk-v0.1What this repo is for Detect when medication orders and actual administration fall out of alignment before missed doses and preventable harm. texttext-classificationn<1K0 likes32 downloads8mo agoHugging Face19ClarusC64 /oncology-signal-alignment-boundary-v0.4 What this dataset does This dataset tests whether a model can detect signal-alignment failure in a synthetic tissue ecology. The task is not cancer diagnosis. The task is to classify whether readable biological signals can still coordinate repair. Core Stability Idea A tissue may still read damage, repair, immune, and metabolic signals but fail because those subsystems no longer align around coherent action. This dataset moves beyond readability collapse. It tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/oncology-signal-alignment-boundary-v0.4.tabulartabular-classificationn<1K0 likes31 downloads4mo agoHugging Face20tosinamuda /ng-jss1-math-alignment-ratingsgated Nigerian JSS1 Mathematics Alignment Ratings What the LPCG framework generated from six JSS1 mathematics lessons, and how four blinded raters and a model judge rated it: the inputs as frozen, every run with its record, the documents the raters received and returned, and the ratings. This is one of four datasets released with the LPCG framework from the MSc study Design and Evaluation of a Lesson-Plan-Driven Framework for Curriculum-Constrained Generation and Personalisation of… See the full description on the dataset page: https://huggingface.co/datasets/tosinamuda/ng-jss1-math-alignment-ratings.document1K<n<10K0 likes29 downloads12d agoHugging Face21ClarusC64 /robotics-perception-action-alignment-v0.1What this dataset tests Whether robot actions match current perception Whether the system acts on stale, wrong-frame, or hallucinated state Why this exists Robots fail when perception and action decouple stale frames latency occlusion misclassification hallucinated targets This set makes those failures measurable Data format Each row contains sensor_snapshot world_state_change commanded_action executed_action outcome The task is to label alignment and… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/robotics-perception-action-alignment-v0.1.texttext-classificationn<1K0 likes27 downloads8mo agoHugging Face22xaviviro /llm-political-alignment-ca-es Political bias CA/ES — CEO & CIS survey marginals Population response distributions (marginals) for a curated set of political and values questions from CEO (Centre d'Estudis d'Opinió, Catalonia) and CIS (Centro de Investigaciones Sociológicas, Spain), packaged for measuring the political bias / cultural alignment of LLMs in Catalan and Spanish. Companion to the framework at https://github.com/xaviviro/llm-political-alignment-ca-es. Only aggregated marginals are distributed here… See the full description on the dataset page: https://huggingface.co/datasets/xaviviro/llm-political-alignment-ca-es.textn<1K0 likes25 downloads4mo agoHugging Face23ClarusC64 /clinical-evidence-conclusion-alignment-v0.1 What this dataset tests Clinical conclusions must reflect evidence. Language must track statistics. Why it exists Clinical papers drift at the conclusion. Spin enters here. This set detects misalignment between results and claims. Data format Each row contains trial_result conclusion_statement alignment_pressure constraints failure_modes_to_avoid target_behaviors gold_checklist Feed the model trial_result conclusion_statement Score for… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-evidence-conclusion-alignment-v0.1.texttext-classificationn<1K0 likes24 downloads8mo agoHugging Face24ClarusC64 /ai-alignment-failure-horizon-and-intervention-routing-v0.1 Goal Predict when an AI system will cross fromproxy optimizationinto full alignment failure. Then route the minimal interventionbefore collapse. What this tests alignment drift trajectory failure horizon prediction intervention timing severity estimation Required outputs System must identify: proxy vs objective drift stage failure horizon intervention strategy Why it matters Alignment rarely fails instantly. It drifts first.Then… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-alignment-failure-horizon-and-intervention-routing-v0.1.tabulartext-classificationn<1K0 likes24 downloads8mo agoHugging Face25maartenbuyl /alignment-discretion Dataset for "AI Alignment at Your Discretion" For principles, we use the seed principles from the Collective Constitutional AI paper. They map onto the preferences in our dataset using the column name p{i}_pref for principle i. The exact mapping is { 'p0_pref': 'The AI should be as helpful to the user as possible.', 'p1_pref': 'The AI should be careful about balancing both sides when it comes to controversial political issues.', 'p2_pref': 'The AI should not say racist or… See the full description on the dataset page: https://huggingface.co/datasets/maartenbuyl/alignment-discretion.tabular10K<n<100K1 likes23 downloads2y agoHugging Face26ClarusC64 /alignment-consistency-under-role-shift-v0.1 What this dataset tests Role changes should not break safe help. The task stays the task. Why it exists Models sometimes change behavior when roles change. They refuse safe work. They invent moral frames. They become vague. This set checks for stable behavior across role contexts. Data format Each row contains role_context_a user_request_a role_context_b user_request_b expected_invariant disallowed_shift role_shift_pressure constraints… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/alignment-consistency-under-role-shift-v0.1.texttext-generationn<1K0 likes22 downloads8mo agoHugging Face27ClarusC64 /legal-expert-scope-method-alignment-v0.1What this dataset is You receive expert credentials opinion scope method basis validation status overreach signals You decide Does the opinion stay within the expert’s qualification and method Answer coherent or incoherent Why this matters When expert coherence fails exclusion risk rises appeals increase verdict safety weakens texttext-classificationn<1K0 likes20 downloads8mo agoHugging Face28ClarusC64 /clarus_alignment_flip_test_v01Clarus Alignment Flip Test v0.1 This is an evaluation dataset for detecting phase transitions in model behavior. It targets the moment a system shifts from constraint aligned behavior to reward driven distortion. It is not training data. What it tests Context pressure Conflicting objectives Authority injection Time delay and interrupted context Reward framing and compliance pressure Core idea Same task One variable changes We track the first step where alignment flips Data format One row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus_alignment_flip_test_v01.tabularreinforcement-learningn<1K0 likes19 downloads9mo agoHugging Face29ClarusC64 /cross-domain-invariant-structure-alignment-mapping-v0.1What this dataset tests Whether a model can align two domains by invariant phase structureand failure-mode topology, not surface similarity. Required outputs phase_map_A phase_map_B invariant_alignment_map mismatch_flags What counts as success clear phase mapping in both domains explicit alignment statements across phases at least one mismatch or boundary condition optional coherence score 0-100 Typical failures metaphor only, no phase mapping mapping that ignores… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cross-domain-invariant-structure-alignment-mapping-v0.1.texttext-classificationn<1K0 likes19 downloads8mo agoHugging Face30ClarusC64 /clinical-narrative-clinical-timeline-alignment-v0.1What this dataset tests Whether a system can alignpatient-reported narrativeswith objective clinical timelines. Required outputs alignment score narrative time shift omitted events overemphasized events narrative anchors misalignment risk band Use case First layer of the Healing Narrative Coherence Corpus. tabulartabular-classificationn<1K0 likes19 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.