agentic-learning-ai-lab/point_maze_medium
PointMaze Medium The PointMaze-Medium dataset from Temporal Straightening for Latent Planning (Wang et al., 2026). It has 4,000 trajectories of 100 steps in the D4RL maze2d-medium-v0 maze, rendered at 224×224. It extends the umaze set in DINO-WM and uses the same file layout. Layout point_maze_medium/ ├── obses/episode_000.pth … episode_3999.pth # uint8 [100, 224, 224, 3] per episode ├── states.pth # float64 [4000, 100, 4]: x, y, vx, vy ├── actions.pth… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/point_maze_medium.
PointMaze Medium
The PointMaze-Medium dataset from Temporal Straightening for Latent Planning (Wang et al., 2026). It has 4,000 trajectories of 100 steps in the D4RL maze2d-medium-v0 maze, rendered at 224×224. It extends the umaze set in DINO-WM and uses the same file layout.
Layout
point_maze_medium/
├── obses/episode_000.pth … episode_3999.pth # uint8 [100, 224, 224, 3] per episode
├── states.pth # float64 [4000, 100, 4]: x, y, vx, vy
├── actions.pth # float64 [4000, 100, 2]: in [-1, 1]
├── seq_lengths.pth # int64 [4000], all 100
├── metadata.pth # generation settings
├── examples/ # mp4 clips of the first ten episodes
└── MANIFEST.sha256 # checksums: sha256sum -c MANIFEST.sha256Frame t and states[:, t] are both taken after step t; t = 0 is the reset. There is no stored split. The paper splits by trajectory at load time, 3,600 train and 400 validation.
Generation
The point-maze environment from the temporal-straightening code, which builds on DINO-WM (d4rl 1.1, mujoco-py 2.1.2.14, MuJoCo 2.1.0, gym 0.23.1). Actions are uniform in [-1, 1]² with numpy seed 0. Episode i starts from env.sample_random_init_goal_states(100 * i).
Loading
import torch
from huggingface_hub import snapshot_download
root = snapshot_download("agentic-learning-ai-lab/point_maze_medium", repo_type="dataset")
states = torch.load(f"{root}/states.pth") # [4000, 100, 4]
actions = torch.load(f"{root}/actions.pth") # [4000, 100, 2]
frames = torch.load(f"{root}/obses/episode_000.pth") # uint8 [100, 224, 224, 3]To use it with the temporal-straightening code, place or symlink the downloaded directory at $DATASET_DIR/point_maze_medium.
License
CC BY 4.0. Please cite the paper below if you use the data.
Citation
@article{wang2026temporal_straightening,
title={Temporal Straightening for Latent Planning},
author={Wang, Ying and Bounou, Oumayma and Zhou, Gaoyue and Balestriero, Randall and Rudner, Tim GJ and LeCun, Yann and Ren, Mengye},
journal={arXiv preprint arXiv:2603.12231},
year={2026}
}
@article{zhou2024dinowm,
title={DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning},
author={Zhou, Gaoyue and Pan, Hengkai and LeCun, Yann and Pinto, Lerrel},
journal={arXiv preprint arXiv:2411.04983},
year={2024}
}
@article{fu2020d4rl,
title={D4RL: Datasets for Deep Data-Driven Reinforcement Learning},
author={Fu, Justin and Kumar, Aviral and Nachum, Ofir and Tucker, George and Levine, Sergey},
journal={arXiv preprint arXiv:2004.07219},
year={2020}
}