datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X2I-in-context-learning
X2I Dataset
Project Page: https://vectorspacelab.github.io/OmniGen/
Github: https://github.com/VectorSpaceLab/OmniGen
Paper: https://arxiv.org/abs/2409.11340
Model: https://huggingface.co/Shitao/OmniGen-v1
To achieve robust multi-task processing capabilities, it is essential to train the OmniGen on large-scale and diverse datasets. However, in the field of unified image generation, a readily available dataset has yet to emerge. For this reason, we have curated a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/yzwang/X2I-in-context-learning.aloha_incontext
aloha_incontext
A Mobile ALOHA robot manipulation dataset for in-context imitation learning. It contains
human-teleoperated demonstrations of pick-and-place, pen uncapping, placing eggs in a
box and closing it, and additional bimanual tasks.
1,328 episodes / 31 task configurations / 587,000 frames, recorded at 50 Hz.
The task configurations are divided into 25 seen configurations (1,318 episodes)
and 6 unseen configurations (10 episodes).
Observations and actions… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/aloha_incontext.word_in_contextDataset homepage:
https://wic-ita.github.io/index.html
gpcv_incontext_benchrecycling-in-common-contextpen_incontextThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 100,
"total_frames": 50000,
"total_tasks": 4,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/pen_incontext.in-context-learning-cosmos3-output
Physical-ICL × Cosmos3 — generated outputs
Video-generation outputs from NVIDIA Cosmos3-Nano (Diffusers Cosmos3OmniPipeline,
image-to-video) on the Physical-ICL dataset (Vincwng/Physical-ICL, subset
physiq_prelim, 66 query samples). This studies physical in-context learning: does
showing a demonstration change how the model continues a query scene?
Total generated: 247 videos across 66 query tasks, in 6 configurations.
Configurations
Every configuration uses the… See the full description on the dataset page: https://huggingface.co/datasets/yqi19/in-context-learning-cosmos3-output.repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Background_INCONTEXTaudio-kw-in-contextin-context-grid-reasoning
In-Context Grid Reasoning (ICGR)
A small, fully synthetic benchmark for demonstration-conditioned rule induction:
each task shows 2–4 (input grid → output grid) support pairs that share one
hidden transformation, and the model must apply the same transformation to a
held-out query input.
It targets the same behaviour probed by recent in-context / latent-reasoning work
on ARC-AGI (e.g. BDH-CQ: In-Context Learning with Recurrent Latent Reasoning,
arXiv:2608.09888), but is… See the full description on the dataset page: https://huggingface.co/datasets/WhySoCodius/in-context-grid-reasoning.wan22-incontext-control-data
Wan2.2 In-Context Control — derived training data
Derived metadata/pose NPZs for the in-context camera + audio control fork of DiffSynth-Studio
(training Wan2.2-TI2V-5B). This repo holds only the small derived files needed to reproduce the
camera and audio (e11h) runs. It does not rehost source videos — those come from the original
datasets linked below.
Companion code: the DiffSynth-Studio fork (see its README for the full reproduction walkthrough).
Files… See the full description on the dataset page: https://huggingface.co/datasets/Haosonnn/wan22-incontext-control-data.geometry3k-in-context-synthesizingThis dataset is used for unsupervised post-training of multi-modal large language models (MLLMs). It contains image-text pairs where the 'problem' field presents a question requiring reasoning and the 'answer' field provides a solution. This data supports the MM-UPT framework detailed in the associated paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
The dataset contains 2101 examples in the… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/geometry3k-in-context-synthesizing.GeoQA-8K-in-context-synthesizing
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
Video-InContext-Editingkeep-reasoning-in-contextincontext_demo4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 16,
"total_frames": 7450,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:16"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/incontext_demo4.incontext_demo2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 1,
"total_frames": 4333,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/incontext_demo2.MMR1-in-context-synthesizingThis dataset is designed for unsupervised post-training of Multi-Modal Large Language Models (MLLMs) focusing on enhancing reasoning capabilities. It contains image-problem-answer triplets, where the problem requires multimodal reasoning to derive the correct answer from the provided image. The dataset is intended for use with the MM-UPT framework described in the accompanying paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/MMR1-in-context-synthesizing.digital-sat-words-in-context-llmGemini-1.5-Pro-In-Context-Invariance-Datasetclimate-2day-inContextQA_incontext_nq_SQuAD_3shot_1docsCaption-Anything-InContextCaption-Anything-InContext is a dataset curated using the model Caption-Pro for improved in-context captioning of images. This model is designed for generating multiple captions for images, ensuring they are contextually accurate.
Required Lib
!pip install -q transformers qwen-vl-utils==0.0.2
Demo with transformers
import os
import gdown
import torch
from transformers import Qwen2VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
from PIL import… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Caption-Anything-InContext.incontext_demoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 1,
"total_frames": 2998,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/incontext_demo.medical-2day-inContextincontext_demo3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "lekiwi",
"total_episodes": 3,
"total_frames": 8587,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/paszea/incontext_demo3.medical-7day-inContextincontext_nq_v2_chunkedword-in-context-definitions
Word-in-Context Definitions Dataset
Training data for fine-tuning LLMs to explain word meanings based on context.
Dataset Description
Each example contains:
word: The target word
context: A sentence containing the word
definition: The meaning of the word in that specific context
Format
Raw format (train.json, val.json, test.json)
{"word": "bright", "context": "a bright sunny day", "definition": "giving out or reflecting much light; shining", "pos":… See the full description on the dataset page: https://huggingface.co/datasets/elcooooo/word-in-context-definitions.
