datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
compositionalityomega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.composition-classifications
NuBerea Composition Classifications
A curated reference set of scholarly-consensus composition history for the biblical corpus: the traditions behind the Old Testament, Deuterocanon, New Testament, and Old Testament Pseudepigrapha, and the source-critical relationships among them (e.g. Documentary Hypothesis strands, Markan priority, canonical collection, translation into the Septuagint). The dataset is a direct transcription of established scholarship — no machine learning or… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/composition-classifications.composition
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig,
Xiang Yue
Carnegie Mellon University, Language Technologies Institute
Does Reinforcement Learning Truly Extend Reasoning?
This work explores the discrepancy in views on RL's effectiveness in extending language models' reasoning abilities. Some characterize RL as a capability refiner, while others see it as inducing new compositional skills. This challenge… See the full description on the dataset page: https://huggingface.co/datasets/Interplay-LM-Reasoning/composition.multi-image-composition-instruction-following
Multi-Image Composition Instruction-Following
A large-scale multimodal dataset for multi-image composition via natural language instruction-following. Each case provides 2-3 input images (characters + scene) along with detailed Chinese instructions to compose them into a single photorealistic output image.
Designed for training and evaluating models on complex image composition tasks that require understanding of character identity preservation, pose generation, scene integration… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multi-image-composition-instruction-following.compositionality_eccv_captioncompositionality_hpsv1Compositional-ARCCompositional-ARC: Assessing Systematic Generalization in Abstract Spatial Reasoning
Philipp Mondorf, Shijia Zhou, Monica Riedler, and Barbara Plank. (2026). Compositional-ARC: Assessing systematic generalization in abstract spatial reasoning. In The Fourteenth International Conference on Learning Representations.
Systematic generalization refers to the capacity to understand and generate novel combinations from known components. Despite recent progress by large language… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/Compositional-ARC.mJev-Compositional-VQA
mJev-Compositional-VQA
English | 简体中文
mJev-Compositional-VQA is a human-reviewed Chinese visual question answering dataset for mJev-style candidate-based evaluation containing 193 questions grounded in 100 images. It evaluates compositional visual grounding: selecting a target by combining visible attributes, spatial relations and reference objects. The release contains natural photographs from COCO, Open Images and Places365, paired with structured Jev questions and… See the full description on the dataset page: https://huggingface.co/datasets/Immortal-Zhang/mJev-Compositional-VQA.gradiend-function-composition
GRADIEND Function Composition Data
Synthetic alias-resolution cloze data used in the GRADIEND/ACTIEND/SAE/CAA comparison (aieng-lab/iend-study).
Usage
from datasets import load_dataset
ds = load_dataset("aieng-lab/gradiend-function-composition", "default", split="train")
Splits: train, validation, test.
Dataset Details
Dataset Description
One variable aliases a variable holding the target value, e.g. m = fork; p = seal; h = m; h =… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/gradiend-function-composition.compositional-safety-folds
Compositional Safety Policy Benchmark — Contrastive Folds
Dataset Summary
This dataset evaluates whether language models apply written safety policies
compositionally, as opposed to responding to lexical features of a request. Each
instance pairs a self-contained policy of seven or eight numbered rules with a
user request, and is labelled with the action the policy requires and the subset
of rules that determine it.
Instances are organised into contrastive folds:… See the full description on the dataset page: https://huggingface.co/datasets/zmsy/compositional-safety-folds.Composition-RL-EVA
Composition-RL
This repository contains the datasets presented in the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that addresses the issue of "too-easy" prompts by automatically composing multiple verifiable problems into a single, more challenging yet still verifiable prompt. RL training on these compositional prompts helps… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Composition-RL-EVA.RL-Compositionality-Stage1-RFT-DataStage 1 RFT data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
mars-chemcam-compositions
Mars ChemCam LIBS Oxide Compositions
Part of the Planetary Science Datasets collection on Hugging Face.
Major oxide compositions of Mars surface rock and soil targets analyzed by the
Chemistry and Camera (ChemCam) Laser-Induced Breakdown Spectroscopy (LIBS)
instrument aboard the Curiosity rover. Currently 30,458 individual
point analyses across 4,184 named targets, spanning sols
0 to 4612.
Dataset description
ChemCam fires a focused laser pulse at rock and soil… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/mars-chemcam-compositions.composition-10B-valvlm-compositionality-embeddings
VLM Compositionality Embeddings
Pre-computed image and text embeddings for the thesis "From Euclidean to Hyperbolic Vision-Language Spaces: A Study of Attribute–Object Compositionality" by Meelad Dashti (Politecnico di Torino & University of Twente, 2026).
Code repository: github.com/MelDashti/hyperbolic-vlm-compositionality
Models
Model
Geometry
Architecture
Training Data
CLIP ViT-L/14
Spherical
ViT-L/14
WIT (400M+ pairs)
DINOv2 ViT-L/14
Spherical
ViT-L/14… See the full description on the dataset page: https://huggingface.co/datasets/Meldashti/vlm-compositionality-embeddings.compositionality_seetruecomposition-10B-testfood-composition-matrix
Food Composition Nutrient Matrix — TKPI 2017 & USDA SR Legacy 2018
This repository contains two food composition datasets reformatted as wide-format nutrient matrices, suitable for a wide range of research tasks including nutrient prediction, food type classification, missing value imputation, dietary analysis, and other machine learning applications on food data. Both datasets share a harmonised set of 18 common nutrients, enabling cross-dataset generalization experiments.… See the full description on the dataset page: https://huggingface.co/datasets/ULM-DS-Lab/food-composition-matrix.CompositionalGSM_augmented
Compositional GSM_augmented
Compositional GSM_augmented is a math instruction dataset, inspired by Not All LLM Reasoners Are Created Equal.
It is based on nvidia/OpenMathInstruct-2 dataset, so you can use this dataset as training dataset.
It is generated using meta-llama/Meta-Llama-3.1-70B-Instruct model by Hyperbloic AI link. (Thanks for free credit!)
Replace the description of the data with the contents in the paper.
Each question in compositional GSM consists of two questions… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/CompositionalGSM_augmented.RL-Compositionality-Stage2-RL-Level8-TestDataStage 2 RL Level 1 to 8 evaluation data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
composition-10B-rlcompositional_celebritiesFood-Composition
Ingredients CSV/Parquet File
Overview
The following data comes from the United States Department of Agriculture’s Food Composition Database. It contains data for various types of food ingredients including the amounts of different vitamins and minerals found in the foods as well as macronutrient percentages. The food covered spans a large variety of foods from butter to Campbell’s soup. Much of the supplementary documenation for each field comes directly from that pages’… See the full description on the dataset page: https://huggingface.co/datasets/hootan09/Food-Composition.glass_alloy_compositionThis is an alloy composition datasetMATH-Composition-199K
Composition-RL
Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
Code | Collection
Composition-RL is a data-efficient approach for Reinforcement Learning with Verifiable Rewards (RLVR). It addresses the issue of "too-easy" prompts (prompts that already achieve a pass rate of 1) by automatically composing multiple verifiable problems into a single, more challenging compositional prompt. This maintains informative training signals and… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-199K.Physics-MATH-Composition-141K
Composition-RL
This repository contains the datasets for the paper Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models.
GitHub | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that combats the growing number of “too-easy” prompts (pass-rate = 1) by automatically composing multiple verifiable problems into a single, harder yet still-verifiable prompt. Across 4B–30B models… See the full description on the dataset page: https://huggingface.co/datasets/xx18/Physics-MATH-Composition-141K.compositionality_image_rewardRL-Compositionality-Stage2-RL-Level1-TrainDataStage 2 RL Level 1 data.
Paper: https://huggingface.co/papers/2509.25123
Code: https://github.com/PRIME-RL/RL-Compositionality
tampa-bay-vacancy-composition-zcta-2020-2024
Tampa Bay vacancy composition by ZCTA, 2020-2024
Author: Richard (Ryszard) Cieplechowicz. DOI: https://doi.org/10.5281/zenodo.23114340 . Study page: https://richardcieplechowicz.com/ . Also on Figshare: https://doi.org/10.6084/m9.figshare.34062054.v1
Richard Cieplechowicz (also legally known as Ryszard Cieplechowicz) · September 29, 2026
A vacancy rate alone does not say why housing units are vacant. In the 132 Tampa Bay ZIP Code Tabulation Areas (ZCTAs) selected for the… See the full description on the dataset page: https://huggingface.co/datasets/richardcieplechowicz/tampa-bay-vacancy-composition-zcta-2020-2024.
