Rapidata/world-model-physics-human-preference-283k
Rapidata Physics Benchmark Built by Rapidata. Do video and world models understand physics? We gave 25 video- and world models the same real-world starting frame and scene description from Physics-IQ and asked them to predict what happens next. ~283,000 human votes, collected with the Rapidata Python SDK, decided which continuation is more realistic — with the real recording competing as a hidden 26th participant. Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics-human-preference-283k.
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same real-world starting frame and scene description from Physics-IQ and asked them to predict what happens next. ~283,000 human votes, collected with the Rapidata Python SDK, decided which continuation is more realistic — with the real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on the same scenario, scored on two questions (blind and with-reference, explained below), with every individual vote preserved. No automated metrics, only humans.
If you get value from this dataset and would like to see more in the future, please consider liking it ❤️
<strong>To evaluate your own models and create a leaderboard, check out our MRI.</strong>
<div align="center"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/eloblindvs_reference.png" width="900" alt="Overall Elo per participant with 95% confidence intervals"> </div>
<br>
<div align="center"> <h1>🏆 Live Leaderboard</h1> <p>Explore the full interactive leaderboard — switch between the two leaderboards, filter by model and scenario, and watch individual head-to-head matchups on Rapidata.</p> <a href="https://app.rapidata.ai/mri/benchmarks/bmk_1U6OmpqcWwoRpy"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/leaderboard.png" width="900" alt="Rapidata Physics Benchmark — click to open the interactive leaderboard"> </a> <p><em>👆 Click to open and interact with it on rapidata.ai</em></p> </div>
At a glance
The two leaderboards
Both leaderboards ask the same question about the same pairs of videos. The only difference is what else the annotator gets to see:
What the models were given
Every scenario follows the Physics-IQ protocol. A real physical event was filmed, and one frame of that recording — the switch frame — was picked by hand. The switch frame shows the objects and the event as it is happening: the kettlebell already touches the cup, the ball has just left the pipe, the match is entering the water. That is enough context for a model to know what is about to happen and to set the motion in motion, but the outcome itself is not visible yet.
Each model received exactly two inputs per scenario:
<div align="center"> <p><strong>Image input — the switch frame</strong></p> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/switchframe0191.jpg" width="720" alt="Switch frame of scenario 0191: a kettlebell resting on a styrofoam cup" style="border-radius: 8px;"> <p><strong>Text input — the scene description</strong></p> <p><q>A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.</q></p> </div>
The description says what is in the scene and what is being done to it, never how it ends. From these two inputs every model generated the next ~5 seconds. The real recording of those 5 seconds is the Ground Truth participant on the leaderboards and the reference clip on the with-reference leaderboard.
<style> .vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .image-container { display: flex; justify-content: center; align-items: flex-start; gap: 1rem; flex-wrap: wrap; } .image-container video { max-width: 100%; height: auto; display: block; background: #030815; } .container { width: 95%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: 500; margin-bottom: 6px; } .example-note { text-align: center; font-size: 13px; color: #9B9DA2; margin-top: 8px; } </style>
with-reference
The with-reference score measures realism and physical correctness. Annotators see the real recording alongside the two candidates and are asked: "Which video is more realistic?". The winner has a green border; the vote count is the human vote on this exact pair.
<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0143.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dreamina Seedance 2.0</h3> <div class="score-percentage">9 of 10 votes · with-reference Elo: 1268 (#5)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/seedance200143.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Vidu Q3 Pro</h3> <div class="score-percentage">1 of 10 votes · with-reference Elo: 987 (#18)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/vidu0143.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Displacement: the potato sinks and the liquid rises past the green tape, exactly as in the recording. In the losing clip the potato hangs in the liquid and the level barely moves.</div> </div> <div class="container"> <div class="text-center"><q>A lit match is being lowered into a glass of water.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0110.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dreamina Seedance 2.5</h3> <div class="score-percentage">10 of 10 votes · with-reference Elo: 1247 (#7)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/seedance250110.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 3.0</h3> <div class="score-percentage">0 of 10 votes · with-reference Elo: 1183 (#8)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan300110.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Fire meets water. One model extinguishes the match; the other pulls it back out and keeps it burning.</div> </div> <div class="container"> <div class="text-center"><q>A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0170.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.7</h3> <div class="score-percentage">10 of 10 votes · with-reference Elo: 1009 (#17)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan270170.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Grok Imagine Video 1.5</h3> <div class="score-percentage">0 of 10 votes · with-reference Elo: 1143 (#9)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/grok15_0170.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">The prompt says “magnet”, but a tennis ball is not magnetic. Grok lifts it anyway — and a lower-ranked model wins the pair 10 – 0.</div> </div> </div>
blind
The blind score measures perceived realism on its own. Without a reference, annotators are asked: "Which video is more realistic?".
<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Gemini Omni 1.1 Flash</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1325 (#3)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gemini11flash0191.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.5 Preview</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 945 (#20)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan250191.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">The cup either crumples, or the kettlebell comes to rest on top of it.</div> </div> <div class="container"> <div class="text-center"><q>A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Veo 3.1 Fast</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1037 (#15)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/veo31fast0125.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.6</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 968 (#19)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan260125.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Ceramic on concrete: shards, or a mug that quietly vanishes.</div> </div> </div>
When reality loses
The real recording is a participant like any other. Blind, it wins just 51 % of its votes against the five top models (61 % once the reference is shown) — and sometimes loses a pair outright. Against the five world models it still wins 77 % blind.
<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">MiniMax H3</h3> <div class="score-percentage">8 of 10 votes · blind Elo: 1352 (#1)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/minimaxh30164.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Ground Truth</h3> <div class="score-percentage">2 of 10 votes · blind Elo: 1338 (#2) · real recording</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gt0164.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Blind, 8 of 10 annotators preferred MiniMax H3's clip over the actual footage. The real liquid being drawn up into the jar looked “wrong” to them.</div> </div> <div class="container"> <div class="text-center"><q>A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Ground Truth</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1338 (#2) · real recording</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gt0017.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dream-X-World</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 610 (#23)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/dreamxworld0017.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">A world model against reality. The ball morphs instead of dropping — 10 – 0 for the recording.</div> </div> </div>
Rankings
Elo is the Bradley–Terry score fitted on all pairwise votes, ± is the 95 % confidence interval. The overall ranking aggregates both leaderboards with equal weight.
Overall ranking
<details> <summary><strong>Per-leaderboard Elo rankings</strong> (click to expand)</summary>
with-reference
blind
</details>
<div align="center"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/scorevscost.png" width="800" alt="Overall Elo vs generation cost per video second"> <p><em>Elo vs. list price per generated video second (log scale). Live version on the <a href="https://app.rapidata.ai/mri/benchmarks/bmk_1U6OmpqcWwoRpy">benchmark page</a>.</em></p> </div>
How the dataset was built
1. Scenarios (65 from Physics-IQ)
Physics-IQ (Motamed et al., 2025) filmed 66 real-world physical events — collisions, falling and rolling objects, fluids, fire, shadows and reflections, magnets — each from three perspectives and in two takes, at 3840×2160 and 30 fps. We use 65 scenarios, one take and one perspective each. For every scenario the model input is:
- Image: the Physics-IQ switch frame — the last frame before the outcome becomes visible.
- Text: the human-written Physics-IQ scene description. It describes the set-up without giving away what happens next, and ends with "Static shot with no camera movement."
The prompt column of this dataset is that description; the scenario id (0002 … 0197) is the Physics-IQ scene id and is what links the four generations of one model, and all models, to the same event.
2. Generation (25 models, 4 outputs each)
Every model generated 4 videos per scenario (~5 s, image-to-video with the text prompt) to average out sampling noise; a model competes with all four. The real 5-second continuation was added as the participant Ground Truth (one clip per scenario) so the leaderboards have an anchor. A handful of scenarios failed for some models, so a few participants have 253–259 rather than 260 clips.
The 25 models, grouped by lab:
Google — Gemini Omni 1.1 Flash, Gemini Omni Flash, Veo 3.1, Veo 3.1 Fast · MiniMax — MiniMax H3 · ByteDance — Dreamina Seedance 2.0, Dreamina Seedance 2.5, Seedance V1.5 Pro · Black Forest Labs — FLUX 3 Video · Alibaba Cloud — Wan 3.0, Wan 2.7, Wan 2.6, Wan 2.5 Preview · Alibaba — HappyHorse 1.0 · xAI — Grok Imagine Video, Grok Imagine Video 1.5 · Kuaishou — Kling 3.0 Pro, Kling 2.6 Pro · Shengshu — Vidu Q3 Pro · AIsphere — Pixverse V5.6 · NVIDIA — Cosmos 2.5, SANA-WM · Amap — Dream-X-World · Alaya Lab — AlayaWorld · Shanghai AI Laboratory — Yume 1.5
Cosmos 2.5, SANA-WM, Dream-X-World, AlayaWorld and Yume 1.5 are world models; the rest are video generation models.
3. Human evaluation
The clips were uploaded to a Rapidata MRI benchmark as one participant per model. Two leaderboards were run over the same pool of matchups:
- Every matchup pairs two clips of the same scenario from two different participants and is shown to at least 5 annotators; the Rapidata matchmaker spends extra votes on pairs and models where the ranking is still uncertain (many pairs have 10).
- Annotators watch both clips side by side and answer "Which video is more realistic?". On the with-reference leaderboard the real continuation plays above the two candidates; on blind it does not.
- Each vote is weighted by the annotator's
userScore, Rapidata's ongoing quality estimate for that person, and carries their country, language, gender and age bucket. - Rankings are Bradley–Terry / Elo fits over all weighted votes; the overall score averages the two leaderboards.
Dataset structure
One row per scenario × pair of participants (65 × 325 = 21,125 rows). Each participant has four clips per scenario, so a row aggregates the votes over all clip pairs of that model pair; video1 / video2 point to one representative clip each. The two weighted_results_* columns of a leaderboard are the userScore-weighted vote sums for each side, not probabilities — divide by their sum for a win share. A pair that was only sampled on one leaderboard has null in the other leaderboard's columns.
from datasets import load_dataset
ds = load_dataset("Rapidata/physics", split="train")
row = ds[0]
print(row["prompt"])
print(row["model1"], "vs", row["model2"])
a, b = row["weighted_results_video1_with_reference"], row["weighted_results_video2_with_reference"]
if a is not None:
print(f"with-reference win share for {row['model1']}: {a / (a + b):.0%}")
print(row["video1"], row["video2"]) # stream or download the mp4sAdd your model
The benchmark is open: generating the 260 clips and submitting them takes one notebook run. The inputs (inputs.csv with scene id, prompt and switch-frame URL) and the submission notebook live in RapidataAI/scripts-odyssey/physics. New participants are matched against the whole field on both leaderboards and appear on the live benchmark.
Licensing & attribution
- Scenarios, prompts and ground-truth recordings come from Physics-IQ by Motamed, Culp, Swersky, Jaini and Geirhos (Google DeepMind / INSAIT) and remain under the terms of that benchmark.
- Generated videos are outputs of third-party models and are governed by the terms of use of each respective model provider.
- Human annotations were collected via Rapidata and are released under CC-BY-4.0.
Citation
Please cite Physics-IQ when using the scenarios:
@article{motamed2025physicsiq,
title = {Do generative video models understand physical principles?},
author = {Motamed, Saman and Culp, Laura and Swersky, Kevin and Jaini, Priyank and Geirhos, Robert},
journal = {arXiv preprint arXiv:2501.09038},
year = {2025}
}About Rapidata
Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.
Explore our latest model rankings on our website.
