datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_market_dynamicsSTUZero-Atari-Dynamics
STUZero Atari Dynamics Dataset
Offline dynamics training datasets collected from trained EfficientZero V2 (EZv2) benchmark models on Atari games. Each game's data is stored in a subfolder named {game}_{steps} indicating the game and the number of training steps of the source checkpoint. While all models were trained for 120K steps, best results in some games were attained at earlier checkpoints. The model with best eval scores was used to curate data for each game.… See the full description on the dataset page: https://huggingface.co/datasets/Shivamkak/STUZero-Atari-Dynamics.humanoid-robots-training-dataset
Dynamic Intelligence — Humanoid Robot Training Dataset
A first-person (egocentric) video dataset of human hand manipulation, designed for training humanoid robot policies via imitation learning. Each episode captures a person performing an everyday household task — folding clothes, moving dishes, opening doors — filmed from a head-mounted iPhone using its built-in LiDAR and depth sensors.
The dataset pairs each video with frame-level 3D hand tracking and camera pose data, giving… See the full description on the dataset page: https://huggingface.co/datasets/DynamicIntelligence/humanoid-robots-training-dataset.enterprise-dynamics-corpus-v1
Enterprise Dynamics Corpus (EDC v1)
Real-data pretraining corpus for enterprise forecasting: commerce and sales, demand and growth, finance and economic activity, IT/cloud operations, infrastructure and industrial asset telemetry. Energy, climate and transport are capped extras, not the core.
Status (2026-09-28): The default commercial_core contains 3.229 GB of verified real training Parquet, 777,504 series records across 18 sources. 46.771 GB remains toward the 50 GB… See the full description on the dataset page: https://huggingface.co/datasets/prabakaranchandran/enterprise-dynamics-corpus-v1.pldr-llm-training-dynamics-data
PLDR-LLM Training Dynamics Data
Reported numerical evidence for Training and Inference Dynamics of PLDR-LLMs:
Row-Map Collapse, Renormalization, and Predictive Reduction, by Burc Gokden.
Monograph: Hugging Face Paper Page.
Scientific code and readers: GitHub repository.
Numerical evidence: Hugging Face dataset.
Book: Power Law Graph Attention and PLDR-LLMs: Mathematical Foundations, Training Dynamics, and Predictive Inference, by Burc Gokden.
Book Companion: Code and edition… See the full description on the dataset page: https://huggingface.co/datasets/fromthesky/pldr-llm-training-dynamics-data.physam4d-dynamics-checkpointsEverMemBench-Dynamic
EverMemBench-Dynamic
A benchmark dataset for evaluating long-term memory capabilities in conversational AI systems. It is part of EverMemBench, the first benchmark designed for long-horizon collaborative memory, introduced in the paper Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues — accepted at KDD 2026 (Oral).
Configurations
This dataset has three configurations (subsets):
dialogues
Multi-turn group dialogues spanning ~250… See the full description on the dataset page: https://huggingface.co/datasets/EverMind-AI/EverMemBench-Dynamic.DynamicMCPBench
DynamicMCPBench
A trace-grounded, effect-scored benchmark for LLM agents on live MCP servers.
Tasks are generated forward: an explorer agent drives real MCP tools until a goal
is reached, the recorded trace is distilled into a TaskSpec, and candidates are graded
on whether they reproduce the effects the trace produced — checkpoints, equivalence
sets, minefields, a partial order — never on matching an answer string or a fixed tool
list. Candidates are evaluated under… See the full description on the dataset page: https://huggingface.co/datasets/TokenWasteGroup/DynamicMCPBench.conversational-dynamics-egocom
Conversational Dynamics — EgoCom
Derived temporal annotations and model-ready training anchors for
conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The
underlying objective is to expose conversational data in a representation
suitable for temporal and action-conditioned models:
state_t + action_t → future conversational state
Contents
Three configurations are provided.… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom.protein_dynamic_properties
Protein Sequences and Dynamic Properties for Training SeqDance and ESMDance
This dataset contains protein sequences, dynamic properties, and feature weights used for training SeqDance and ESMDance, two protein language models designed to learn protein dynamic properties.
training_test_data_sequence.csv
This file contains 64,403 protein sequences with associated metadata for training and testing (excluding dynamicPDB).
Columns:
name – Unique identifier for each… See the full description on the dataset page: https://huggingface.co/datasets/ChaoHou/protein_dynamic_properties.DynamicRAG_Training_Data_150kDynamicVLA-Task-Examples
DynamicVLA Simulation Task Examples
打开 MP4 任务浏览页
DOM has 3 task families, not 3 task_index values: 140,932 structured instruction IDs and 207,306 episodes. This subset selects 6 demonstrations per family (18 total), preserving actual task_index/episode_index and structured instructions. All 3 published cameras are single-arm cameras: opposite, wrist, side, not top/left/right arms.
所有示例是完整 H.264 MP4。单臂相机保留官方名称:opst_cam / wrist_cam / side_cam。每条有同步合并视频、各路原视频与官方动作/状态 Parquet。… See the full description on the dataset page: https://huggingface.co/datasets/Travor278/DynamicVLA-Task-Examples.Semantic-Flow-Dynamics-SFD
Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus
TL;DR: 614 Chinese-language formalized social-science concepts across 25
papers, UUID-linked with typed derivation relations (derives_from,
leads_to, falsified_by, …) — usable for knowledge-graph construction,
RAG over structured theory, or as a Chinese formal-reasoning corpus.
Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com
What This Dataset Is
This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.groundtruth-dynamic-benchmarking
Groundtruth Dynamic Benchmarking — Geology
Question sets and grading rubrics for evaluating LLMs on real-world geological
reasoning. Every question is authored from a real source corpus, and every
claim in the grading key carries an evidence locator back to that corpus —
nothing is synthetic. Licensing/redistribution status varies by corpus — see
License.
This dataset holds the questions, grading rubrics, and source corpora.
Running an evaluation (generating answers from a model… See the full description on the dataset page: https://huggingface.co/datasets/EigenformAI/groundtruth-dynamic-benchmarking.conversational-dynamics-egocom-12.5hz
Conversational Dynamics — EgoCom (12.5 Hz)
Derived temporal annotations and model-ready training anchors for
conversational-dynamics and turn-taking research, generated from EgoCom.
This dataset is produced by the conversational-dynamics-data pipeline. The
underlying objective is to expose conversational data in a representation
suitable for temporal and action-conditioned models:
state_t + action_t → future conversational state
12.5 Hz release
The same pipeline… See the full description on the dataset page: https://huggingface.co/datasets/batgre/conversational-dynamics-egocom-12.5hz.SpeakerVerification_Aishell1Train
Dataset Card for "SpeakerVerification_AISHELL1Train"
More Information needed
DynamicvlmSpeechTextMatching_LibrispeechTrainClean360
Dataset Card for "speechTextMatching_LibrispeechTrainClean360"
More Information needed
Disney-Theme-Park-Queue-Dynamics
🎢 Disney World Queue Dynamics
A Comprehensive EDA & Strategic Analysis
Author: Matan Zigelman • University: Reichman University • Date: March 2026
📋 Project Introduction: Disney Theme Park Queue Dynamics
This project analyzes a numeric-heavy operational dataset from a major Disney theme park, sourced from kaggle and containing approximately 3,757,301 records. The dataset is primarily driven by time-based and operational metrics ( e.g.… See the full description on the dataset page: https://huggingface.co/datasets/matanzig/Disney-Theme-Park-Queue-Dynamics.EnhancementDetection_LibrittsTrainClean360Wham
Dataset Card for "EnhancementDetection_LibrittsTrainClean360Wham"
More Information needed
World-Embedding-Dynamics-Retrieval
World Embedding Dynamics Retrieval
World Embedding Dynamics Retrieval is the dynamics retrieval split of the
World Embedding Benchmark. It contains
1,900 simulation videos from 19 physics families. The family shards are loaded
together as the default configuration.
Usage
from datasets import load_dataset
dataset = load_dataset(
"World-Representation-Lab/World-Embedding-Dynamics-Retrieval",
split="test",
)
Fields
query_id: unique… See the full description on the dataset page: https://huggingface.co/datasets/World-Representation-Lab/World-Embedding-Dynamics-Retrieval.SpeechTextMatching_Tedlium2Train
Dataset Card for "SpeechTextMatching_TEDLIUM2Train"
More Information needed
SpeechDetection_Aishell1Train
Dataset Card for "SpeechDetection_AISHELL1Train"
More Information needed
fresh-swe-pro-dynamic
Dataset Card
Dataset Description
Fresh SWE-Pro Dynamic is a source-verified, Docker-executable benchmark for software-engineering agents. The public release contains five recent repository tasks with pinned base commits, problem statements, gold patches, regression tests, Docker images, source provenance, and validation evidence.
Task: software-engineering agent evaluation and patch generation
Language: English
Public size: 5 instances
Release: 2026.09
Source… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/fresh-swe-pro-dynamic.DiSH-Bench-dynamic-v1
DiSH-Bench dynamic v1
2,000 physical dynamic scenes, 4,532 original four-channel FOA recordings and 48,396 test questions. Easy/Hard each contain 1,000 scenes. The 500 matched groups each contain easy_human, easy_robot, hard_joint_a and hard_joint_b; room, initial positions, speech crops, utterance timing and playback roles are shared within a group. All source/receiver RIR knots and spatial task outputs use the same 0.2-second scene clock.
This is a separate dynamic release… See the full description on the dataset page: https://huggingface.co/datasets/Heiheihaha17/DiSH-Bench-dynamic-v1.dynamic_robot_bench_dr_scripted_10k
dynamic_robot_bench_dr_scripted_10k
10,000 scripted-expert demonstrations across all 100 dynamic task families of
dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka
Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset:
100 episodes per family, success-filtered, language-prompted per episode.
Collection configuration (identical for every family)
Scripted expert with per-step auto-derived speed caps, recorded as… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_10k.ChordClassification_AcousticGuitarAndPiano
Dataset Card for "chord_classification_acoustic_guitar_and_piano"
More Information needed
SpeakerVerification_Tedlium2Train
Dataset Card for "SpeakerVerification_TEDLIUM2Train"
More Information needed
NoiseSNRLevelPredictionGaussian_VoxcelebMusan
Dataset Card for "NoiseSNRLevelPredictiongaussian_VoxcelebMusan"
More Information needed
RespiratorySoundClassification_ICBHI2017
