datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
composition
Difficulty Split (Zero Context Medium)
This dataset mirrors the local layout used in training:
train/*.jsonl
val_id/*.jsonl
val_ood/*.jsonl
Each JSONL row contains fields like problem, question, and solution (the latter includes an Answer: segment near the end).
Load with datasets (streaming)
from datasets import load_dataset
repo = "goodevening/composition"
train = load_dataset(
"json",
data_files={"train": f"hf://datasets/{repo}/train/*.jsonl"}… See the full description on the dataset page: https://huggingface.co/datasets/goodevening/composition.compositionalityomega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.composition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.MaternKernel_compositionalityYou can load the dataset as follows:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="shc443/MaternKernel_compositionality", repo_type="dataset")
For more information regarding data generating process, please refer to our paper or github page
wellborn-indian-food-compositionThis package provides detailed nutritional values for 542 key foods in India, based on direct measurements across six regions. Data was obtained from the book Indian Food Composition Tables 2017, published by the National Institute of Nutrition, Hyderabad.
▌
📦 JSR,
📦 NPM,
📦 Corpus,
📰 Docs,
🌐 Website.
import * as ifct2017 from "jsr:@nodef/ifct2017";
await ifct2017.loadCompositions();
await ifct2017.loadColumns();
await ifct2017.loadIntakes();
// Load corpus first… See the full description on the dataset page: https://huggingface.co/datasets/Uvathe/wellborn-indian-food-composition.galahad-deconf-compositional
Deconfounded compositional set (RoboCasa)
Part of the Galahad release · Project page · Code + battery + generator
504 episodes; attribute conjunction (size × colour) with single-attribute foils: each foil matches the target on exactly one attribute, so neither word alone identifies it.
LeRobot v2.1 format. Generated by generator/collect_rc_compositional.py in the release repo; the generator produces the exam (the battery) and the medicine (the training set) from the same code.… See the full description on the dataset page: https://huggingface.co/datasets/phi-monster/galahad-deconf-compositional.VL-Compositionality-Benchmarks
VL-Compositionality-Benchmarks
This repository contains a collection of evaluation benchmarks used to assess the compositional understanding of Vision-Language Models (VLMs), as presented in the paper Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality.
Official GitHub Repository: hiker-lw/MACCO
Dataset Summary
These benchmarks are designed to evaluate how well models like CLIP capture object relations… See the full description on the dataset page: https://huggingface.co/datasets/hiker-lw/VL-Compositionality-Benchmarks.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.compositionality_eccv_captioncompositionality_hpsv1Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations.
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.mJev-Compositional-VQA
mJev-Compositional-VQA
English | 简体中文
mJev-Compositional-VQA is a human-reviewed Chinese visual question answering dataset for mJev-style candidate-based evaluation containing 193 questions grounded in 100 images. It evaluates compositional visual grounding: selecting a target by combining visible attributes, spatial relations and reference objects. The release contains natural photographs from COCO, Open Images and Places365, paired with structured Jev questions and… See the full description on the dataset page: https://huggingface.co/datasets/Immortal-Zhang/mJev-Compositional-VQA.gradiend-function-composition
GRADIEND Function Composition Data
Synthetic alias-resolution cloze data used in the GRADIEND/ACTIEND/SAE/CAA comparison (aieng-lab/iend-study).
Usage
from datasets import load_dataset
ds = load_dataset("aieng-lab/gradiend-function-composition", "default", split="train")
Splits: train, validation, test.
Dataset Details
Dataset Description
One variable aliases a variable holding the target value, e.g. m = fork; p = seal; h = m; h =… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/gradiend-function-composition.thinkworld_final_oven_robot_human_composition_v1
ThinkWorld oven composition prompt v1
This is an additive metadata derivation of
chyun/thinkworld_final_oven_robot_human@a0c9dcca7a9e2ba394c306bcdf6e7161b712dbc0. Video bytes, timing, robot
state/action labels, supervision masks, camera layout, and source provenance are
unchanged. Human episodes retain their original global composite prompt exactly.
The ordinary LeRobot task resolves from task_index to the episode-global
instruction. atomic_task_index retains the current skill… See the full description on the dataset page: https://huggingface.co/datasets/chyun/thinkworld_final_oven_robot_human_composition_v1.eval2_compositional_augmentedcompositional_causal_reasoning
– 3k+ Hugging Face downloads –
https://jmaasch.github.io/ccr/
Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring these behaviors requires principled
evaluation methods. Maasch et al. (2025) consider both behaviors simultaneously, under
the umbrella of compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate
through graphs. CCR.GB applies the… See the full description on the dataset page: https://huggingface.co/datasets/jmaasch/compositional_causal_reasoning.compositional-safety-folds
Compositional Safety Policy Benchmark — Contrastive Folds
Dataset Summary
This dataset evaluates whether language models apply written safety policies
compositionally, as opposed to responding to lexical features of a request. Each
instance pairs a self-contained policy of seven or eight numbered rules with a
user request, and is labelled with the action the policy requires and the subset
of rules that determine it.
Instances are organised into contrastive folds:… See the full description on the dataset page: https://huggingface.co/datasets/zmsy/compositional-safety-folds.causal-pybullet-dual-franka-composition-v1
Causal PyBullet Dual-Franka Composition v1
Dataset summary
causal_pybullet_dual_franka_composition_v1 is a deterministic synthetic
dual-arm manipulation dataset generated in PyBullet with two Franka Panda
robots. It targets long-horizon compositional manipulation, causal
generalization, direct pick-and-place, and cross-table handover.
The release contains successful closed-loop physical rollouts rather than
kinematic label-only demonstrations. Object transfers are… See the full description on the dataset page: https://huggingface.co/datasets/zhb10086/causal-pybullet-dual-franka-composition-v1.Composition-RL-EVA
Composition-RL
This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Composition-RL-EVA.RL-Compositionality-Stage1-RFT-DataStage 1 RFT data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
mars-chemcam-compositions
Mars ChemCam LIBS Oxide Compositions
Part of the Planetary Science Datasets collection on Hugging Face.
Major oxide compositions of Mars surface rock and soil targets analyzed by the
Chemistry and Camera (ChemCam) Laser-Induced Breakdown Spectroscopy (LIBS)
instrument aboard the Curiosity rover. Currently 30,458 individual
point analyses across 4,184 named targets, spanning sols
0 to 4612.
Dataset description
ChemCam fires a focused laser pulse at rock and soil… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/mars-chemcam-compositions.composition-10B-valvlm-compositionality-embeddings
VLM Compositionality Embeddings
Pre-computed image and text embeddings for the thesis "From Euclidean to Hyperbolic Vision-Language Spaces: A Study of Attribute–Object Compositionality" by Meelad Dashti (Politecnico di Torino & University of Twente, 2026).
Code repository: github.com/MelDashti/hyperbolic-vlm-compositionality
Models
Model
Geometry
Architecture
Training Data
CLIP ViT-L/14
Spherical
ViT-L/14
WIT (400M+ pairs)
DINOv2 ViT-L/14
Spherical
ViT-L/14… See the full description on the dataset page: https://huggingface.co/datasets/Meldashti/vlm-compositionality-embeddings.compositionality_seetruecomposition-10B-testfood-composition-matrix
Food Composition Nutrient Matrix — TKPI 2017 & USDA SR Legacy 2018
This repository contains two food composition datasets reformatted as wide-format nutrient matrices, suitable for a wide range of research tasks including nutrient prediction, food type classification, missing value imputation, dietary analysis, and other machine learning applications on food data. Both datasets share a harmonised set of 18 common nutrients, enabling cross-dataset generalization experiments.… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/food-composition-matrix.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.RL-Compositionality-Stage2-RL-Level8-TestDataStage 2 RL Level 1 to 8 evaluation data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
