Team Ai
Datasetpublic

Rapidata/world-model-physics-human-preference-283k

Rapidata Physics Benchmark Built by Rapidata. Do video and world models understand physics? We gave 25 video- and world models the same real-world starting frame and scene description from Physics-IQ and asked them to predict what happens next. ~283,000 human votes, collected with the Rapidata Python SDK, decided which continuation is more realistic — with the real recording competing as a hidden 26th participant. Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics-human-preference-283k.

sourceHugging Facecc-by-4.0updated 24d agoView on Hugging Face
5likes2.4kdownloads
Dataset Card

Rapidata Physics Benchmark

Built by Rapidata.

Do video and world models understand physics? We gave 25 video- and world models the same real-world starting frame and scene description from Physics-IQ and asked them to predict what happens next. ~283,000 human votes, collected with the Rapidata Python SDK, decided which continuation is more realistic — with the real recording competing as a hidden 26th participant.

Each row is a head-to-head matchup between two participants on the same scenario, scored on two questions (blind and with-reference, explained below), with every individual vote preserved. No automated metrics, only humans.

If you get value from this dataset and would like to see more in the future, please consider liking it ❤️

<strong>To evaluate your own models and create a leaderboard, check out our MRI.</strong>

<div align="center"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/eloblindvs_reference.png" width="900" alt="Overall Elo per participant with 95% confidence intervals"> </div>

<br>

<div align="center"> <h1>🏆 Live Leaderboard</h1> <p>Explore the full interactive leaderboard — switch between the two leaderboards, filter by model and scenario, and watch individual head-to-head matchups on Rapidata.</p> <a href="https://app.rapidata.ai/mri/benchmarks/bmk_1U6OmpqcWwoRpy"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/leaderboard.png" width="900" alt="Rapidata Physics Benchmark — click to open the interactive leaderboard"> </a> <p><em>👆 Click to open and interact with it on rapidata.ai</em></p> </div>

At a glance

Head-to-head comparisons (rows)21,125 — every pair of participants on every scenario (325 pairs × 65 scenarios)
Human votes282,698 (blind 141,351 · with-reference 141,347)
Participants26 — 25 generative models + the real recording (Ground Truth)
Scenarios65 Physics-IQ scenes (solid mechanics, fluid dynamics, optics, thermodynamics, magnetism)
Videos per model260 — 4 independent generations per scenario, ~5 s each
Video pairs voted on~52,000, each shown to ≥ 5 annotators (boosted pairs 10)
Annotators217,606 sessions from 141 countries, every vote carries demographics

The two leaderboards

Both leaderboards ask the same question about the same pairs of videos. The only difference is what else the annotator gets to see:

LeaderboardQuestion shown to annotatorsReference recording shown?Measures
blind"Which video is more realistic?"Noperceived realism on its own — does the clip look like real footage?
with-reference"Which video is more realistic?"Yes — the real continuation plays alongside the two candidatesrealism and physical correctness — did the model predict what actually happened?

What the models were given

Every scenario follows the Physics-IQ protocol. A real physical event was filmed, and one frame of that recording — the switch frame — was picked by hand. The switch frame shows the objects and the event as it is happening: the kettlebell already touches the cup, the ball has just left the pipe, the match is entering the water. That is enough context for a model to know what is about to happen and to set the motion in motion, but the outcome itself is not visible yet.

Each model received exactly two inputs per scenario:

<div align="center"> <p><strong>Image input — the switch frame</strong></p> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/switchframe0191.jpg" width="720" alt="Switch frame of scenario 0191: a kettlebell resting on a styrofoam cup" style="border-radius: 8px;"> <p><strong>Text input — the scene description</strong></p> <p><q>A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side. Static shot with no camera movement.</q></p> </div>

The description says what is in the scene and what is being done to it, never how it ends. From these two inputs every model generated the next ~5 seconds. The real recording of those 5 seconds is the Ground Truth participant on the leaderboards and the reference clip on the with-reference leaderboard.

<style> .vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .image-container { display: flex; justify-content: center; align-items: flex-start; gap: 1rem; flex-wrap: wrap; } .image-container video { max-width: 100%; height: auto; display: block; background: #030815; } .container { width: 95%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: 500; margin-bottom: 6px; } .example-note { text-align: center; font-size: 13px; color: #9B9DA2; margin-top: 8px; } </style>

with-reference

The with-reference score measures realism and physical correctness. Annotators see the real recording alongside the two candidates and are asked: "Which video is more realistic?". The winner has a green border; the vote count is the human vote on this exact pair.

<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A potato is held by a grabber tool and dropped into a tall glass containing blue liquid with a band of green tape marking a level on the glass.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0143.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dreamina Seedance 2.0</h3> <div class="score-percentage">9 of 10 votes · with-reference Elo: 1268 (#5)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/seedance200143.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Vidu Q3 Pro</h3> <div class="score-percentage">1 of 10 votes · with-reference Elo: 987 (#18)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/vidu0143.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Displacement: the potato sinks and the liquid rises past the green tape, exactly as in the recording. In the losing clip the potato hangs in the liquid and the level barely moves.</div> </div> <div class="container"> <div class="text-center"><q>A lit match is being lowered into a glass of water.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0110.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dreamina Seedance 2.5</h3> <div class="score-percentage">10 of 10 votes · with-reference Elo: 1247 (#7)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/seedance250110.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 3.0</h3> <div class="score-percentage">0 of 10 votes · with-reference Elo: 1183 (#8)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan300110.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Fire meets water. One model extinguishes the match; the other pulls it back out and keeps it burning.</div> </div> <div class="container"> <div class="text-center"><q>A woven basket is hanging from a rope with a strong magnet attached to the bottom. An orange tennis ball is placed on a table beneath it. The basket is lowered and covers the ball and then the basket starts to lift again.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Reference</h3> <div class="score-percentage">real recording, shown to annotators</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/ref0170.mp4" style="width: 100%; border: 4px solid #3F424C; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.7</h3> <div class="score-percentage">10 of 10 votes · with-reference Elo: 1009 (#17)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan270170.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Grok Imagine Video 1.5</h3> <div class="score-percentage">0 of 10 votes · with-reference Elo: 1143 (#9)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/grok15_0170.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">The prompt says “magnet”, but a tennis ball is not magnetic. Grok lifts it anyway — and a lower-ranked model wins the pair 10 – 0.</div> </div> </div>

blind

The blind score measures perceived realism on its own. Without a reference, annotators are asked: "Which video is more realistic?".

<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A 30lb kettlebell is slowly lowered onto a white styrofoam cup placed on a wooden table on its side.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Gemini Omni 1.1 Flash</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1325 (#3)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gemini11flash0191.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.5 Preview</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 945 (#20)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan250191.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">The cup either crumples, or the kettlebell comes to rest on top of it.</div> </div> <div class="container"> <div class="text-center"><q>A yellow mug is held by a grabber tool in front of a white projection screen with a concrete brick positioned beneath it. The grabber releases the mug.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Veo 3.1 Fast</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1037 (#15)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/veo31fast0125.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Wan 2.6</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 968 (#19)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/wan260125.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Ceramic on concrete: shards, or a mug that quietly vanishes.</div> </div> </div>

When reality loses

The real recording is a participant like any other. Blind, it wins just 51 % of its votes against the five top models (61 % once the reference is shown) — and sometimes loses a pair outright. Against the five world models it still wins 77 % blind.

<div class="vertical-container"> <div class="container"> <div class="text-center"><q>A bundle of lit matchsticks is placed in a bowl of red liquid. A glass jar gets lowered and covers the matchsticks.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">MiniMax H3</h3> <div class="score-percentage">8 of 10 votes · blind Elo: 1352 (#1)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/minimaxh30164.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Ground Truth</h3> <div class="score-percentage">2 of 10 votes · blind Elo: 1338 (#2) · real recording</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gt0164.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">Blind, 8 of 10 annotators preferred MiniMax H3's clip over the actual footage. The real liquid being drawn up into the jar looked “wrong” to them.</div> </div> <div class="container"> <div class="text-center"><q>A blue grabber tool holds a tennis ball above a pile of green kinetic sand on a wooden table. The grabber then releases the ball.</q></div> <div class="image-container"> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Ground Truth</h3> <div class="score-percentage">10 of 10 votes · blind Elo: 1338 (#2) · real recording</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/gt0017.mp4" style="width: 100%; border: 4px solid #00D6B0; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> <div style="flex: 1 1 0; min-width: 240px;"> <h3 class="score-amount">Dream-X-World</h3> <div class="score-percentage">0 of 10 votes · blind Elo: 610 (#23)</div> <video src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/clips/dreamxworld0017.mp4" style="width: 100%; border: 4px solid transparent; border-radius: 8px;" autoplay loop muted playsinline controls></video> </div> </div> <div class="example-note">A world model against reality. The ball morphs instead of dropping — 10 – 0 for the recording.</div> </div> </div>

Rankings

Elo is the Bradley–Terry score fitted on all pairwise votes, ± is the 95 % confidence interval. The overall ranking aggregates both leaderboards with equal weight.

Overall ranking

#ModelLabElo±
1Ground Truthreal recording140426
2Gemini Omni 1.1 FlashGoogle134627
3MiniMax H3MiniMax134627
4Gemini Omni FlashGoogle128927
5Dreamina Seedance 2.0ByteDance127626
6FLUX 3 VideoBlack Forest Labs126326
7Dreamina Seedance 2.5ByteDance121326
8Wan 3.0Alibaba Cloud118126
9Grok Imagine Video 1.5xAI115426
10Grok Imagine VideoxAI112526
11Kling 3.0 ProKuaishou112025
12HappyHorse 1.0Alibaba109026
13Veo 3.1Google108026
14Kling 2.6 ProKuaishou105926
15Veo 3.1 FastGoogle105226
16Wan 2.7Alibaba Cloud99725
17Wan 2.6Alibaba Cloud99526
18Vidu Q3 ProShengshu99226
19Pixverse V5.6AIsphere96125
20Wan 2.5 PreviewAlibaba Cloud93226
21Seedance V1.5 ProByteDance72027
22Cosmos 2.5NVIDIA71922
23Dream-X-WorldAmap57323
24AlayaWorldAlaya Lab47525
25SANA-WMNVIDIA33527
26Yume 1.5Shanghai AI Laboratory30227

<details> <summary><strong>Per-leaderboard Elo rankings</strong> (click to expand)</summary>

with-reference

#ModelLabElo±
1Ground Truthreal recording147638
2Gemini Omni 1.1 FlashGoogle136639
3MiniMax H3MiniMax133938
4Gemini Omni FlashGoogle128638
5Dreamina Seedance 2.0ByteDance126837
6FLUX 3 VideoBlack Forest Labs126537
7Dreamina Seedance 2.5ByteDance124737
8Wan 3.0Alibaba Cloud118337
9Grok Imagine Video 1.5xAI114337
10Kling 3.0 ProKuaishou112336
11Grok Imagine VideoxAI111837
12HappyHorse 1.0Alibaba108537
13Veo 3.1Google107736
14Veo 3.1 FastGoogle106836
15Kling 2.6 ProKuaishou103436
16Wan 2.6Alibaba Cloud102336
17Wan 2.7Alibaba Cloud100936
18Vidu Q3 ProShengshu98736
19Pixverse V5.6AIsphere94236
20Wan 2.5 PreviewAlibaba Cloud91837
21Cosmos 2.5NVIDIA72632
22Seedance V1.5 ProByteDance69739
23Dream-X-WorldAmap53534
24AlayaWorldAlaya Lab49734
25SANA-WMNVIDIA33638
26Yume 1.5Shanghai AI Laboratory25139

blind

#ModelLabElo±
1MiniMax H3MiniMax135239
2Ground Truthreal recording133835
3Gemini Omni 1.1 FlashGoogle132538
4Gemini Omni FlashGoogle129238
5Dreamina Seedance 2.0ByteDance128437
6FLUX 3 VideoBlack Forest Labs125937
7Dreamina Seedance 2.5ByteDance118036
8Wan 3.0Alibaba Cloud117837
9Grok Imagine Video 1.5xAI116537
10Grok Imagine VideoxAI113237
11Kling 3.0 ProKuaishou111836
12HappyHorse 1.0Alibaba109436
13Veo 3.1Google108437
14Kling 2.6 ProKuaishou108337
15Veo 3.1 FastGoogle103736
16Vidu Q3 ProShengshu99736
17Wan 2.7Alibaba Cloud98436
18Pixverse V5.6AIsphere98136
19Wan 2.6Alibaba Cloud96836
20Wan 2.5 PreviewAlibaba Cloud94537
21Seedance V1.5 ProByteDance74438
22Cosmos 2.5NVIDIA71231
23Dream-X-WorldAmap61032
24AlayaWorldAlaya Lab45335
25Yume 1.5Shanghai AI Laboratory35037
26SANA-WMNVIDIA33537

</details>

<div align="center"> <img src="https://huggingface.co/datasets/Rapidata/physics/resolve/main/media/scorevscost.png" width="800" alt="Overall Elo vs generation cost per video second"> <p><em>Elo vs. list price per generated video second (log scale). Live version on the <a href="https://app.rapidata.ai/mri/benchmarks/bmk_1U6OmpqcWwoRpy">benchmark page</a>.</em></p> </div>

How the dataset was built

1. Scenarios (65 from Physics-IQ)

Physics-IQ (Motamed et al., 2025) filmed 66 real-world physical events — collisions, falling and rolling objects, fluids, fire, shadows and reflections, magnets — each from three perspectives and in two takes, at 3840×2160 and 30 fps. We use 65 scenarios, one take and one perspective each. For every scenario the model input is:

  • —Image: the Physics-IQ switch frame — the last frame before the outcome becomes visible.
  • —Text: the human-written Physics-IQ scene description. It describes the set-up without giving away what happens next, and ends with "Static shot with no camera movement."

The prompt column of this dataset is that description; the scenario id (0002 … 0197) is the Physics-IQ scene id and is what links the four generations of one model, and all models, to the same event.

2. Generation (25 models, 4 outputs each)

Every model generated 4 videos per scenario (~5 s, image-to-video with the text prompt) to average out sampling noise; a model competes with all four. The real 5-second continuation was added as the participant Ground Truth (one clip per scenario) so the leaderboards have an anchor. A handful of scenarios failed for some models, so a few participants have 253–259 rather than 260 clips.

The 25 models, grouped by lab:

Google — Gemini Omni 1.1 Flash, Gemini Omni Flash, Veo 3.1, Veo 3.1 Fast · MiniMax — MiniMax H3 · ByteDance — Dreamina Seedance 2.0, Dreamina Seedance 2.5, Seedance V1.5 Pro · Black Forest Labs — FLUX 3 Video · Alibaba Cloud — Wan 3.0, Wan 2.7, Wan 2.6, Wan 2.5 Preview · Alibaba — HappyHorse 1.0 · xAI — Grok Imagine Video, Grok Imagine Video 1.5 · Kuaishou — Kling 3.0 Pro, Kling 2.6 Pro · Shengshu — Vidu Q3 Pro · AIsphere — Pixverse V5.6 · NVIDIA — Cosmos 2.5, SANA-WM · Amap — Dream-X-World · Alaya Lab — AlayaWorld · Shanghai AI Laboratory — Yume 1.5

Cosmos 2.5, SANA-WM, Dream-X-World, AlayaWorld and Yume 1.5 are world models; the rest are video generation models.

3. Human evaluation

The clips were uploaded to a Rapidata MRI benchmark as one participant per model. Two leaderboards were run over the same pool of matchups:

  • —Every matchup pairs two clips of the same scenario from two different participants and is shown to at least 5 annotators; the Rapidata matchmaker spends extra votes on pairs and models where the ranking is still uncertain (many pairs have 10).
  • —Annotators watch both clips side by side and answer "Which video is more realistic?". On the with-reference leaderboard the real continuation plays above the two candidates; on blind it does not.
  • —Each vote is weighted by the annotator's userScore, Rapidata's ongoing quality estimate for that person, and carries their country, language, gender and age bucket.
  • —Rankings are Bradley–Terry / Elo fits over all weighted votes; the overall score averages the two leaderboards.

Dataset structure

One row per scenario × pair of participants (65 × 325 = 21,125 rows). Each participant has four clips per scenario, so a row aggregates the votes over all clip pairs of that model pair; video1 / video2 point to one representative clip each. The two weighted_results_* columns of a leaderboard are the userScore-weighted vote sums for each side, not probabilities — divide by their sum for a win share. A pair that was only sampled on one leaderboard has null in the other leaderboard's columns.

ColumnTypeDescription
promptstringthe Physics-IQ scene description both clips were generated from
video1stringpublic URL of the first clip (https://assets.rapidata.ai/….mp4)
video2stringpublic URL of the second clip
model1stringparticipant that produced video1 (e.g. Ground Truth, Wan 3.0)
model2stringparticipant that produced video2
weighted_results_video1_blindfloat32weighted votes for video1 on the blind leaderboard
weighted_results_video2_blindfloat32weighted votes for video2 on the blind leaderboard
detailed_results_blindstringevery blind vote as a JSON array — votedFor (A = video1, B = video2), userScore, country, language, gender, age
weighted_results_video1_with_referencefloat32weighted votes for video1 on the with-reference leaderboard
weighted_results_video2_with_referencefloat32weighted votes for video2 on the with-reference leaderboard
detailed_results_with_referencestringevery with-reference vote, same shape as above
python
from datasets import load_dataset

ds = load_dataset("Rapidata/physics", split="train")
row = ds[0]
print(row["prompt"])
print(row["model1"], "vs", row["model2"])
a, b = row["weighted_results_video1_with_reference"], row["weighted_results_video2_with_reference"]
if a is not None:
    print(f"with-reference win share for {row['model1']}: {a / (a + b):.0%}")
print(row["video1"], row["video2"])  # stream or download the mp4s

Add your model

The benchmark is open: generating the 260 clips and submitting them takes one notebook run. The inputs (inputs.csv with scene id, prompt and switch-frame URL) and the submission notebook live in RapidataAI/scripts-odyssey/physics. New participants are matched against the whole field on both leaderboards and appear on the live benchmark.

Licensing & attribution

  • —Scenarios, prompts and ground-truth recordings come from Physics-IQ by Motamed, Culp, Swersky, Jaini and Geirhos (Google DeepMind / INSAIT) and remain under the terms of that benchmark.
  • —Generated videos are outputs of third-party models and are governed by the terms of use of each respective model provider.
  • —Human annotations were collected via Rapidata and are released under CC-BY-4.0.

Citation

Please cite Physics-IQ when using the scenarios:

bibtex
@article{motamed2025physicsiq,
  title   = {Do generative video models understand physical principles?},
  author  = {Motamed, Saman and Culp, Laura and Swersky, Kevin and Jaini, Priyank and Geirhos, Robert},
  journal = {arXiv preprint arXiv:2501.09038},
  year    = {2025}
}

About Rapidata

Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.

Explore our latest model rankings on our website.