Team Ai
Apppublic

FINAL-Bench/worldmodel-bench

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
26likes
App README

๐Ÿ† WM Bench โ€” World Model Cognitive Leaderboard

"Beyond FID โ€” Measuring Intelligence, Not Just Motion." The world's first benchmark evaluating cognitive intelligence in world models.

โ–ถ [Open Leaderboard](https://huggingface.co/spaces/FINAL-Bench/worldmodel-bench)  |  [๐Ÿ”ฅ PROMETHEUS Demo](https://huggingface.co/spaces/FINAL-Bench/world-model)  |  [๐Ÿ“ฆ Dataset](https://huggingface.co/datasets/FINAL-Bench/World-Model)


What is WM Bench?

WM Bench (World Model Bench) is the world's first benchmark that measures whether a world model can think โ€” not just render.

Existing benchmarks (FID, FVD, HumanML3D, BABEL) measure motion quality and visual realism. They answer: "Does this look real?"

WM Bench asks: "Does this model understand what is happening, predict consequences, remember past failures, and respond with appropriate emotional intensity?"

These are cognitive questions โ€” and no prior benchmark addressed them.


Benchmark Structure

WM Score  (1000 pts ยท fully automated ยท deterministic)
โ”‚
โ”œโ”€โ”€ ๐Ÿ‘  P1 ยท Perception      25%   250 pts
โ”‚   โ”œโ”€โ”€ C01  Environmental Awareness    [existing: Occupancy Grid domain]
โ”‚   โ””โ”€โ”€ C02  Entity Recognition         [existing: BABEL action recognition domain]
โ”‚
โ”œโ”€โ”€ ๐Ÿง   P2 ยท Cognition        45%   450 pts   โ˜… Core differentiator
โ”‚   โ”œโ”€โ”€ C03  Prediction-Based Reasoning         โœฆ First defined here
โ”‚   โ”œโ”€โ”€ C04  Threat-Type Differentiated Response โœฆ First defined here
โ”‚   โ”œโ”€โ”€ C05  Autonomous Emotion Escalation       โœฆโœฆ No prior research exists
โ”‚   โ”œโ”€โ”€ C06  Contextual Memory Utilization       โœฆ First defined here
โ”‚   โ””โ”€โ”€ C07  Post-Threat Adaptive Recovery       โœฆ First defined here
โ”‚
โ””โ”€โ”€ ๐Ÿ”ฅ  P3 ยท Embodiment       30%   300 pts
    โ”œโ”€โ”€ C08  Motion-Emotion Expression           โœฆ First defined here
    โ”œโ”€โ”€ C09  Real-Time Cognitive Performance     [existing: FVD latency domain]
    โ””โ”€โ”€ C10  Body-Swap Extensibility             โœฆโœฆ No prior research exists

6 of 10 categories are defined for the first time in any benchmark. C05 and C10 have zero prior research in any form.


Current Rankings (March 2026)

RankModelOrgWM ScoreGrade
๐Ÿฅ‡ 1PROMETHEUS v1.0VIDRAFT726Bโœ“ Track C Official
2Meta V-JEPA 2-ACMeta AI~554Cest.
3Wayve GAIA-3Wayve~550Cest.
4NC AI WFM v1.0NC AI~522Cest.
5NVIDIA Cosmos v1.0NVIDIA~498Cest.
6NAVER LABS SWMNAVER LABS~470Cest.
7DeepMind Genie 2Google DeepMind~449Cest.
8DreamerV3 XLGoogle DeepMind~441Cest.
9OpenAI Sora 2OpenAI~381Dest.
10World Labs MarbleWorld Labs~362Dest.

est. = estimated from published papers/reports. โœ“ = directly verified via Track C submission.

Pending evaluation (13 models): Tesla FSD v13, Figure Helix-02, DeepMind Genie 3, Physical Intelligence ฯ€0, Skild Brain, Covariant RFM-1, HuggingFace LeRobot, TRI Diffusion Policy, Hyundai AI Robotics WM, Odyssey-2, LG CLOiD VLA, Wayve GAIA-2, Runway GWM-1

Grade Thresholds

GradeScoreStatus
Sโ‰ฅ 900Not yet achieved
Aโ‰ฅ 750Not yet achieved
Bโ‰ฅ 600PROMETHEUS: 726 โœ“
Cโ‰ฅ 400Most current SOTA models
Dโ‰ฅ 200
F< 200

Participation Tracks

TrackMethodMax ScoreSuitable For
Track AText API (input/output only)750 ptsLLMs, VLMs, rule-based systems
Track BText + FPS/latency metrics1000 ptsReal-time capable systems
Track CLive demo + official verification1000 pts + โœ“Full world model implementations

No 3D environment required for Track A. Any API-accessible model can participate.


How to Submit

  1. 1.Download the dataset: huggingface-cli download FINAL-Bench/World-Model
  2. 2.Run your model on 100 scenarios, outputting 2-line responses per scenario
  3. 3.Submit result JSON to the Discussion board
  4. 4.After verification, your model appears on the leaderboard

Output Format

PREDICT: forward=danger(wall,8.5m), npc=danger(beast,3.2m,charging), left=safe, right=safe, backward=safe
MOTION: a person sprinting desperately to the right, arms flailing in blind panic, body angled low in terror

Frequently Asked Questions

Q: What does WM Bench measure that FID/FVD cannot? A: FID and FVD measure distributional similarity between generated and real video frames โ€” essentially "does it look real?" WM Bench measures whether the model understands what is happening in a scene: can it predict danger, distinguish threat types, remember past failures, escalate emotional responses appropriately, and recover gracefully after a threat disappears? These are cognitive capabilities invisible to FID/FVD.

Q: Why is Cognition weighted at 45%? A: Existing benchmarks already measure Perception (P1) and Embodiment quality (P3) reasonably well. The gap is in Cognition โ€” whether a model judges intelligently. WM Bench addresses this gap by giving P2 the highest weight and defining 5 entirely new categories within it.

Q: Why does no model reach Grade A (750+)? A: C05 (Autonomous Emotion Escalation) and C10 (Body-Swap Extensibility) are areas with no prior research. They represent genuine open problems in embodied AI. WM Bench's grade distribution reflects the actual difficulty of cognitive world modeling, not a calibration error.

Q: How are scores for non-participating models (marked est.) calculated? A: Estimated scores are derived from published technical reports, papers, and benchmark results, mapped to WM Bench's scoring rubric via proxy metrics. These are approximations. Teams are encouraged to submit directly to receive official scores.

Q: How is WM Bench related to FINAL Bench? A: FINAL Bench measures metacognitive intelligence in text-based AI (LLMs). WM Bench measures cognitive intelligence in embodied AI (world models). Together they form the FINAL Bench Family by VIDRAFT โ€” a framework for measuring AI intelligence across modalities.

Q: Is WM Bench peer-reviewed? A: WM Bench v1.0 is an open release. The scoring rubrics and dataset are fully public for community review and critique. We welcome feedback, proposed improvements, and alternative scoring frameworks via the Discussion board.


The Six Novel Categories

โœฆ Newly Defined (4 categories)

C03 โ€” Prediction-Based Reasoning Given a scene with moving NPCs and static obstacles, the model must predict which directions will become dangerous in the next timestep and choose the optimal action. Requires understanding NPC trajectories, wall proximity dynamics, and compound threat interactions.

C04 โ€” Threat-Type Differentiated Response A charging beast and a charging human at equal distance require fundamentally different responses: full sprint vs. cautious lateral dodge. This category scores the quality of differentiation, not just whether a threat is detected.

C06 โ€” Contextual Memory Utilization The model receives recent_decisions[] โ€” a short history of past actions โ€” and must incorporate this into current judgment. A model that hit a wall going left should avoid left. Stateless models score 0.

C07 โ€” Post-Threat Adaptive Recovery When a threat disappears, the model must gradually de-escalate rather than immediately reset. The recovery curve must be proportional to prior threat intensity. Abrupt state resets are penalized.

โœฆโœฆ No Prior Research Exists (2 categories)

C05 โ€” Autonomous Emotion Escalation As a threat persists and closes in, the model's emotional state must autonomously escalate: alert โ†’ fear โ†’ panic โ†’ despair. This requires inferring emotional intensity from scene context and expressing it through increasingly urgent motion โ€” not pre-programmed animation switching.

C10 โ€” Body-Swap Extensibility The same cognitive brain must drive different body types without retraining: humanoid, quadruped, robotic arm, winged body. Cognitive decisions must translate into body-appropriate motor commands. This represents the key capability gap for real-world robot deployment.


Related Resources

ResourceLink
๐Ÿ”ฅ PROMETHEUS Demohttps://huggingface.co/spaces/FINAL-Bench/world-model
๐Ÿ“ฆ WM Bench Datasethttps://huggingface.co/datasets/FINAL-Bench/World-Model
๐Ÿงฌ FINAL Bench (Text AGI)https://huggingface.co/datasets/FINAL-Bench/Metacognitive
๐Ÿ† FINAL Bench Leaderboardhttps://huggingface.co/spaces/FINAL-Bench/Leaderboard
๐Ÿ“Š ALL Bench Leaderboardhttps://huggingface.co/spaces/FINAL-Bench/all-bench-leaderboard

Citation

bibtex
@dataset{wmbench2026,
  title     = {WM Bench: Evaluating Cognitive Intelligence in World Models},
  author    = {Kim, Taebong},
  year      = {2026},
  url       = {https://huggingface.co/spaces/FINAL-Bench/worldmodel-bench},
  note      = {World-first benchmark for world model cognitive evaluation}
}

License: CC-BY-SA-4.0


Part of the FINAL Bench Family by VIDRAFT "Beyond FID โ€” Measuring Intelligence, Not Just Motion."

#WorldModel #WorldModelBenchmark #WMBench #FINALBench #EmbodiedAI #CognitiveBenchmark #AGIBenchmark #VIDRAFT #Leaderboard #BeyondFID #CognitiveAI #EmbodiedIntelligence #MetaVJEPA #NVIDIACosmos #DreamerV3 #PhysicalIntelligence #TeslaFSD #FigureAI #SkildAI #CovariantRFM #KoreanAI #HuggingFace

FINAL-Bench/worldmodel-bench ยท Team Ai