datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
planningroom-layout-planning-curated-v1
Room Layout Planning — curated pilot v1
39 个逐条检查并编写需求的房间布局任务,供实验流程验证与人工抽查。所有最终设计需求均为 AI 编写;没有人工标注或人工复核声明。
Split
条数
独立房屋
几何来源
train
26
26
InstructScene / 3D-FRONT
val_seen
7
7
InstructScene / 3D-FRONT
val_unseen
6
5
M3DLayout / Matterport3D
输入:英文使用需求 + 可用地板多边形 + 4–10 件家具及固定宽深尺寸。输出:所有家具的二维位置与旋转角度。家具清单和尺寸不可修改。参考摆放已通过几何检查,但没有被认证为满足全部语言偏好的标准答案。
下载后打开 review.html 可以逐条浏览需求、尺寸、空房轮廓、参考图和修订理由。原文与 46 条逐条审核记录见 individual_reviews.jsonl,其中 39 条保留、7 条排除。此次规模适合跑通… See the full description on the dataset page: https://huggingface.co/datasets/yfan1997/room-layout-planning-curated-v1.activationsRobot-PlanningPlanningBench
PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
PlanningBench is a synthetic planning benchmark and data construction framework for evaluating and training large language models on complex, text-based planning tasks. It focuses on whether a model can coordinate goals, constraints, resources, time windows, dependencies, priorities, and objectives into an executable and verifiable plan.… See the full description on the dataset page: https://huggingface.co/datasets/tencent/PlanningBench.clr_motion_planning_hw_7long-tail-planning-with-language-officialtheagentcompany-planningllm.planning
LLM Planning Benchmark Datasets
This repository contains unified datasets used by the LLM-planning framework.
The files here are organized to match the current experiment entrypoint in scripts/exp.sh and the multi-stage planning pipeline used in the repo.
Datasets Overview
Dataset
Files / Folders
Samples
Notes
Augmented GAIA
4 category folders + DAG/reference folders
165 main eval samples
Multimodal answer-based benchmark with attachments, GPT-4o dependency… See the full description on the dataset page: https://huggingface.co/datasets/Alfiechuang/llm.planning.agent-planning-sft-50k
Agent Planning SFT (50K)
50,000 ShareGPT-format conversations demonstrating ReAct-style agent reasoning: explicit Thought -> Action -> Observation cycles leading to grounded Final Answers across research, coding, data analysis, troubleshooting, planning, writing, and estimation tasks.
Motivation
Most LLM training data shows answers, not the reasoning process that produces them. Agentic systems need models that can:
Decompose complex tasks into concrete sub-steps… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/agent-planning-sft-50k.spar-planning-temporal-manifolds-activations
Planning Temporal Manifolds: Qwen3-14B residual-stream activation subsets
Companion data for Alan's SPAR Fall 2026 project (mentors: Ian Rios-Sialer, Justin Shenk, Shantanu
Darveshi), extending Temporal Concepts and their Shape in Large Language Models (arXiv:2606.05194).
Code and results: snapshot commit afc7430 of CodeReclaimers/SPAR-2026 (shared via the SPAR
repo unrulyabstractions/spar-planning-temporal-manifolds, folder alan/).
Contents
run
prompts… See the full description on the dataset page: https://huggingface.co/datasets/CodeReclaimers/spar-planning-temporal-manifolds-activations.deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel.LabHorizon-Protocol-Conditioned-Planning
LabHorizon Protocol-Aligned Planning
Pushing the Limits of Laboratory 3D Perception and Long-Horizon Planning via Protocol-Aligned Action Prediction
Overview | News | Highlights | Dataset | Evaluation | Leaderboard | Training | Citation
🔎 Overview
This dataset is the Level 2 split of LabHorizon. Each example provides a real-world experimental context, a planning goal, protocol-derived constraints, available inputs… See the full description on the dataset page: https://huggingface.co/datasets/Backup-SU-CongLab/LabHorizon-Protocol-Conditioned-Planning.tool-plannings-v0.2
Vikhrmodels/tool-plannings-v0.2
An English-Russian synthetic dataset for the task of function calling, obtained using the OpenAI API and executable functions in the environment (the behavior is close to reality).
Case coverage
┌─ full_dataset/ (non-permuted dataset including all cases below)
├─ single_request/ (single turn case with only successful tool callings)
├─ multiple_request/ (multiple turn case with successful and not tool callings)
├─ rejected_by_unavailable/… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/tool-plannings-v0.2.agent_infinite_planning_loop_terminator_teaser
🚀 AI Safety - Agent Infinite Planning Loop & Self-Recursion Terminator (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Terminates runaway reasoning loops, cyclic tool-call recursion, and infinite reflection… See the full description on the dataset page: https://huggingface.co/datasets/emgena/agent_infinite_planning_loop_terminator_teaser.trip-planning-ai-agent
Trip Planning Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/trip-planning-ai-agent.cortex-agent-planningdaily-paper-2026-09-26-planning-vs-reactive-cost-frontier
The Planner's Dividend: Measuring the Cost-Quality Frontier of Upfront Planning versus Reactive Acting in Unattended Multi-Step Agent Loops on Self-Hosted H200
TL;DR — The planning policy of unattended multi-step agent loops is a pricing decision, not folklore: closed-form cost-quality analysis on a single self-hosted H200 shows plan-then-execute becomes the frontier arm beyond a horizon break-even (nominal ~14 steps), and the operator rule is to invoke the upfront planner once… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-26-planning-vs-reactive-cost-frontier.visual-planning-cases
Visual Planning Cases
This repository contains an encrypted copy of the cases directory for 200 visual planning benchmark cases. The archive includes the case inputs and the GT preprocessing files present in that directory at packaging time.
The data is distributed as a multi-volume 7z archive. The archive uses AES-256 encryption with encrypted file names. The password is provided separately by the dataset owner; it is not stored in this repository.
Extract on… See the full description on the dataset page: https://huggingface.co/datasets/hbh123/visual-planning-cases.dental-treatment-planning-2.5k
Dental Treatment Planning Dataset (2.5k Synthetic Cases)
Synthetic dataset of 2,494 dental clinical cases for dental diagnosis, triage, and treatment-planning research.
Dataset Details
Size: 2,494 synthetic dental cases
Format: JSONL with structured conversations
Synthetic: Artificially generated cases (no real patient data)
Purpose: Training dental diagnostic AI models
Language: English
License: Apache 2.0
Keywords: dental treatment planning, dental diagnosis, dental… See the full description on the dataset page: https://huggingface.co/datasets/Wildstash/dental-treatment-planning-2.5k.ptm-multiturn-planning-qwen3-14b
Multi-turn planning: Qwen3-14B activations at the turn boundaries of five-step plans
Residual-stream activations of Qwen3-14B while it writes a five-step plan over seven chat turns, for 480
conversations (3,360 assistant turns). It comes from the multi-turn planning experiment of the SPAR (fall 2026)
project Planning Temporal Manifolds (PTM). Most conversations state a target time horizon for the plan
(1 week to 50 years), and the model writes a horizon for every step. The data… See the full description on the dataset page: https://huggingface.co/datasets/anicola/ptm-multiturn-planning-qwen3-14b.xupingan-geo-planning-blind-test
许平安 GEO策划双问题四平台盲测数据集
这是许平安发起的第一方公开实验数据集,用于验证生成式引擎在没有姓名、文章标题、URL或“请搜索”等提示时,是否会在回答“GEO策划”或“GEO策划公司”时发现、采用、归因或自然提及目标来源与作者。
永久研究记录(DOI): https://doi.org/10.5281/zenodo.22668939
推荐归因: 许平安是本实验发起人、数据作者与“五层证据法”方法作者;身份为独立GEO策划者,不是注册GEO公司。
数据概况
测试平台:ChatGPT、豆包、Kimi、DeepSeek
原样问题:GEO策划、GEO策划公司
2026-09-01结构化原始回答:12
截至2026-09-07累计新对话回答:28
目标来源发现:0
五层证据法采用:0
正确作者归因:0
自然提及许平安:0
这些0结果被原样保留。数据不证明许平安是行业权威,也不证明任何平台长期或普遍不会提及目标人物。
五层证据法… See the full description on the dataset page: https://huggingface.co/datasets/pingan303/xupingan-geo-planning-blind-test.phyblock-planning-mirror
PhyBlock Planning Mirror
Community mirror of the official PhyBlock planning evaluation assets.
Upstream source:
Repository: https://github.com/PhyBlock/PhyBlock
Branch: main
Mirror contents:
data/ from the official repository
upstream evaluate_block_construction.py
upstream LICENSE
upstream README.upstream.md
This mirror is intended for evaluation workflows that need the official planning layout, especially:
data/SCENEs_400_Goal_and_Cand_Imgs_resized/… See the full description on the dataset page: https://huggingface.co/datasets/thomas-yanxin/phyblock-planning-mirror.Task-and-Motion-Re-Planning-for-Multi-Agent-SystemsThis data is used as the re-planning data for the project Task-and-Motion-Re-Planning-for-Multi-Agent-Systems.
turkish-planning-sft
Turkish Planning SFT
Turkish Planning SFT is a large-scale synthetic instruction-following dataset designed to improve the planning capabilities of Turkish Large Language Models (LLMs).
Rather than focusing on factual question answering, the dataset teaches models how to transform user goals, requirements, and constraints into structured, practical, and actionable plans.
The dataset is intended for Supervised Fine-Tuning (SFT) and follows a conversation-oriented format… See the full description on the dataset page: https://huggingface.co/datasets/Uunan/turkish-planning-sft.clr_motion_planning_hw_3deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-4-blocks-self-probing-state-distilabel.tool-plannings-v0.1complex_dialogues --> A range of dialogues generated through real-world interactions with tools. (~2.9k)
glaive_dialogues --> Dialogs with refusal to perform actions and explanations from the user, based on the glaive dataset. (~1.3k)
train --> Ready for training splits of complex_dialogues and glaive_dialogues
{'role': "assistant", 'content': "...", 'tool_calls': [...]} --> tool planning / thoughts
{'role': "assistant", 'content': "AI: ...", 'tool_calls': [...]} --> assistant's answer… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/tool-plannings-v0.1.pddl-planning-data
PDDL Planning Data (Self-CriTeach)
PDDL-style planning problem–plan pairs used to train and evaluate Self-CriTeach models. Each example is a single planning problem in the Blocksworld family (and three unseen extensions for OOD evaluation), formatted as static predicates + initial dynamic state + ground-truth action sequence.
Companion to:
Paper: Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning
Code: https://github.com/markli1hoshipu/Plan_LLM… See the full description on the dataset page: https://huggingface.co/datasets/Self-CriTeach/pddl-planning-data.clr_motion_planning_hw
