datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
interactive-stage2-30k-v1InteractiveRadiologyTutor_datssetsSWE-Bench-Pro-interactive-issue-qa
SWE-Bench Pro / Interactive / Issue+QA
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). This dataset contains the QA-augmented trajectories: verifiable questions about repository behavior inserted before, between, or after the split issue turns. Every question ships with its hidden reference answer, an executable golden proof script, and the creation-time proof-execution report.
The plain… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue-qa.qualcomm-interactive-cooking-dataset-ego-mistake-corrections
Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark
Description
This dataset contains cooking videos with timestamped instruction and feedback for task guidance.
Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp.
Dataset Details
Release files:
annotations/annotations.json
videos/*.MP4
Release statistics:
Total videos: 40
Total released annotations: 1,597
Text type… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.SWE-Bench-Pro-interactive-issue
SWE-Bench Pro / Interactive / Issue
Companion data release for the anonymous paper "Opinion: Coding-Agent Benchmarks Should Match Their Users' Task Flows" (SWE-TaskFlow). SWE-TaskFlow transforms an issue-derived benchmark into replayable multi-turn trajectories while preserving the original tasks and tests. This dataset contains the issue-solving prompt sequences (no QA turns) over the 701 SWE-Bench Pro tasks that admit a three-part decomposition (out of the 731 public tasks).… See the full description on the dataset page: https://huggingface.co/datasets/Anonym01048/SWE-Bench-Pro-interactive-issue.Chinese_interactive_novels_3k
中文互动小说结构化语料
This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total.
All contents are parsed from certain online sources.
Usage
This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents.
Each novel is a dict structured as follows:
class Novel:
book_title: str
book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.qualcomm-interactive-cooking-dataset
Qualcomm Interactive Cooking Dataset
Description
The Qualcomm Interactive Cooking Dataset is designed to evaluate the ability of multi-modal LLMs to provide step-by-step instructions, focusing on the cooking domain.
Dataset Details
The Qualcomm Interactive Cooking Dataset includes step-by-step instructions and feedback pairs. The videos are from the CaptainCook4D dataset - licensed under Apache 2.0.
Dataset Collection Process
The text annotations and… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset.interactive-sweDataset Summary
Interactive SWE-bench is a dataset developed by CMU Language Technologies Institute (LTI) that contains 500 verified samples from the SWE-bench test set. This dataset is an enhanced version of the original SWE-bench dataset, featuring both the original detailed GitHub issues and their simplified, focused versions.
The dataset collects 500 test Issue-Pull Request pairs from popular Python repositories. Each entry includes both the original detailed issue description and a… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/interactive-swe.qualcomm-interactive-cooking-dataset-counterfactual-mistakes
Qualcomm Interactive Cooking Dataset: Ego Counterfactual Mistakes
Description
This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps.
Dataset Details
Files:
annotations.json
Release statistics:
Total rows: 25,087
Unique videos (dataset + video_id): 1,110
Rows by source dataset:
CaptainCook4D: 4,969
Ego4D: 13,847
Ego-Exo4D: 6… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-counterfactual-mistakes.paper_universe_interactive
Paper Universe Interactive Graph
Small static-viewer payload for the Research Library paper universe.
This dataset is intentionally separate from PeytonT/paper_graph. It contains browser-friendly interactive assets needed by the static WebGL/WASM app. Parquet is the preferred payload format; JSON is retained as a compatibility fallback:
parquet/interactive/papers_50000.parquet
parquet/interactive/papers_200000.parquet
parquet/interactive/papers_all.parquet… See the full description on the dataset page: https://huggingface.co/datasets/PeytonT/paper_universe_interactive.Xianxia-Cultivation-System-Interactive-Sandbox-System-ExampleInteractive-PEDES-v1
📁 Interactive-PEDES Dataset
This folder contains data for the Interactive-PEDES benchmark used in the paper: LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-Identification" (ICML 2025)
📦 Structure
.
├── CUHK-PEDES
│ ├── caption_all.json
│ ├── imgs/
│ │ ├── cam_a/
│ │ ├── cam_b/
│ │ ├── CUHK01/
│ │ ├── CUHK03/
│ │ ├── Market/
│ │ ├── test_query/
│ │ └── train_query/
├── ICFG-PEDES
│ ├── ICFG-PEDES.json
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/XLearning-SCU/Interactive-PEDES-v1.interactive-sports-nhl
interactive-sports: NHL research database
One SQLite file, 2.46 GB, covering 2010-10-07 to 2026-06-14: 16 tables and
~23M rows of NHL box scores, play-by-play, shifts, and 126,967 dated news notes.
It is the database the agents in
interactive_sports query.
Agents never read it directly. The harness builds cutoff-scoped views over it,
filtered to game_date <= as_of_date, with every player and team replaced by an
opaque P#### / T#### token minted fresh per run.
Use… See the full description on the dataset page: https://huggingface.co/datasets/gilberty005/interactive-sports-nhl.DeepSeek-R1-Distill-Data-5krepo_universe_interactive
Repo Universe Interactive Graph
Small static-viewer payload for the Research Library repository universe.
This dataset is intentionally separate from the full repo graph export. It contains browser-friendly graph assets for the static WebGL/WASM app. Parquet is the preferred payload format; JSON/JSONL is retained as a compatibility fallback:
parquet/interactive/lod_10000.parquet
parquet/interactive/lod_50000.parquet
parquet/interactive/lod_200000.parquet… See the full description on the dataset page: https://huggingface.co/datasets/PeytonT/repo_universe_interactive.humanoid-interactive-dialogue-states
Humanoid Interactive Dialogue States
A state-based interactive dialogue dataset for humanoid robots.
Interactive_Benchmarks
Interactive Benchmarks (IB)
Interactive Benchmarks for Evaluating Interactive Reasoning and Agent Capabilities
Usage
from datasets import load_dataset
dataset = load_dataset("interactivebench/Interactive_Benchmarks")
Citation
If you use the Interactive_Benchmarks dataset in your research, please consider citing it as follows:
@misc{interactivebench,
title={Interactive Benchmarks for Evaluating Interactive Reasoning and Agent… See the full description on the dataset page: https://huggingface.co/datasets/interactivebench/Interactive_Benchmarks.3d-interactive-assestsinteractive-fiction-knowledgeReasoning-While-Asking-SFT-Datasethumanoid-interactive-session-logs
Humanoid Interactive Session Logs
High-quality session-level interaction logs for humanoid AI agents.
LLMBind-GPT-Interactive-DataChatMed_Consult_Dataset_Interactiveinteractive-mapunreal-interactive-devhumanoid-interactive-control-panel-logs
Interactive Control Panel Logs
High-fidelity logs representing professional humanoid control interfaces.
strl-main-ec-playwright_interactive-gc-claude_client_strl_summary-mc-claude_agent_sonnet-r0strl-main-ec-playwright_interactive-gc-claude_client_strl_dplm-mc-claude_agent_sonnet-rb-r0Interactive_BenchmarksNekoQA-Interactive
NekoQA-Interactive
Dataset Description
The NekoQA Interactive Dataset is designed to provide a rich and engaging conversational experience between users and their virtual feline companions. This dataset features well-structured JSON samples that include diverse instructions and outputs, showcasing a variety of emotional expressions and interactive scenarios. The purpose of this dataset is to enhance the emotional connection and entertainment value during interactions… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/NekoQA-Interactive.
