datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aloha_incontext
aloha_incontext
A Mobile ALOHA robot manipulation dataset for in-context imitation learning. It contains
human-teleoperated demonstrations of pick-and-place, pen uncapping, placing eggs in a
box and closing it, and additional bimanual tasks.
1,328 episodes / 31 task configurations / 587,000 frames, recorded at 50 Hz.
The task configurations are divided into 25 seen configurations (1,318 episodes)
and 6 unseen configurations (10 episodes).
Observations and actions… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/aloha_incontext.recycling-in-common-contextpen_incontextThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 100,
"total_frames": 50000,
"total_tasks": 4,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vo2yager/pen_incontext.in-context-learning-cosmos3-output
Physical-ICL × Cosmos3 — generated outputs
Video-generation outputs from NVIDIA Cosmos3-Nano (Diffusers Cosmos3OmniPipeline,
image-to-video) on the Physical-ICL dataset (Vincwng/Physical-ICL, subset
physiq_prelim, 66 query samples). This studies physical in-context learning: does
showing a demonstration change how the model continues a query scene?
Total generated: 247 videos across 66 query tasks, in 6 configurations.
Configurations
Every configuration uses the… See the full description on the dataset page: https://huggingface.co/datasets/yqi19/in-context-learning-cosmos3-output.Background_INCONTEXTgeometry3k-in-context-synthesizingThis dataset is used for unsupervised post-training of multi-modal large language models (MLLMs). It contains image-text pairs where the 'problem' field presents a question requiring reasoning and the 'answer' field provides a solution. This data supports the MM-UPT framework detailed in the associated paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
The dataset contains 2101 examples in the… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/geometry3k-in-context-synthesizing.GeoQA-8K-in-context-synthesizing
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO (arXiv:2505.22453)
MMR1-in-context-synthesizingThis dataset is designed for unsupervised post-training of Multi-Modal Large Language Models (MLLMs) focusing on enhancing reasoning capabilities. It contains image-problem-answer triplets, where the problem requires multimodal reasoning to derive the correct answer from the provided image. The dataset is intended for use with the MM-UPT framework described in the accompanying paper.
🐙 GitHub Repo: waltonfuture/MM-UPT
📜 Paper (arXiv): Unsupervised Post-Training for Multi-Modal LLM Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/MMR1-in-context-synthesizing.incontext-best50
