Team Ai
Datasetpublic

VCG-EAI/TraceSpatial-Trace

TraceSpatial-Trace Project · Paper · Code TraceSpatial-Trace is the RGB-and-QA tracing subset of TraceSpatial, introduced by the RoboTracer project. It supports learning to translate language instructions into spatial waypoints for object manipulation and robot end-effector motion. This release contains 517,215 referenced RGB images and 3,623,880 question–answer pairs, extracted from 1,067,822 original conversation records across AgiBot, CA-1M, DROID, RoboTwin, and ScanNet. It… See the full description on the dataset page: https://huggingface.co/datasets/VCG-EAI/TraceSpatial-Trace.

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
4likes4.9kdownloads
Dataset Card

TraceSpatial-Trace

Project · Paper · Code

TraceSpatial-Trace is the RGB-and-QA tracing subset of TraceSpatial, introduced by the RoboTracer project. It supports learning to translate language instructions into spatial waypoints for object manipulation and robot end-effector motion.

This release contains 517,215 referenced RGB images and 3,623,880 question–answer pairs, extracted from 1,067,822 original conversation records across AgiBot, CA-1M, DROID, RoboTwin, and ScanNet. It combines reconstructed indoor scenes, real-world manipulation data, and simulated robot interactions. These are the counts of this released subset, rather than the approximately 30 million QA pairs in the broader TraceSpatial collection described in the paper.

Tasks

  • —2D tracing: predict an ordered sequence of image-plane waypoints (u, v).
  • —3D tracing: predict waypoints with depth (u, v, d).
  • —2D-to-3D lifting: add depth to 2D waypoints supplied in the question.

Image-plane coordinates are normalized to [0, 1000], and depth is expressed in meters. The three-component answer uses image coordinates plus depth, not Cartesian XYZ. This release provides RGB images and QA text only; depth maps and precomputed spatial features are not included.

Data format

Each Parquet row contains one image and one QA pair:

image, question, answer, type, source, task, id, qa_index, record_index.

type preserves the original object / end_effector label and remains null when absent. task is derived from each QA pair. Original question and answer text is unchanged. Use (source, record_index, qa_index) to identify a row; group by (source, record_index) to recover its conversation.

Usage

python
from datasets import load_dataset

ds = load_dataset(
    "leeibo/TraceSpatial-Trace", "all", split="train", streaming=True
)
example = next(iter(ds))
print(example["question"])
print(example["answer"])

Source-specific configurations are also available: agibot, ca1m, droid, robotwin, and scannet. The train label is a packaging split; this release does not define a held-out evaluation set.

License

This dataset release is distributed under the Apache License 2.0. <img src="https://api.visitorbadge.io/api/combined?path=https%3A%2F%2Fzhoues.github.io&labelColor=%232ccce4&countColor=%230158f9" alt="visitor badge" style="display: none;" /> <img src="https://api.visitorbadge.io/api/combined?path=https%3A%2F%2Fanjingkun.github.io&labelColor=%232ccce4&countColor=%230158f9" alt="visitor badge" style="display: none;" />

Reference

bibtex
@article{zhou2025robotracer,
    title={RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics},
    author={Zhou, Enshen and Chi, Cheng and Li, Yibo and An, Jingkun and Zhang, Jiayuan and Rong, Shanyu and Han, Yi and Ji, Yuheng and Liu, Mengzhen and Wang, Pengwei and others},
    journal={arXiv preprint arXiv:2512.13660},
    year={2025}
}