Team Ai
Datasetpublic

Keh0t0/scene-mem-benchmark

scene-mem-benchmark A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work in a house for several minutes, then is asked to retrieve an object it has to remember — one that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the job it was interrupted in, remembering how far it had got — or (routine) to put a new object away where this household keeps that kind of thing, a rule it was never told and can only… See the full description on the dataset page: https://huggingface.co/datasets/Keh0t0/scene-mem-benchmark.

sourceHugging Facemitupdated 19d agoView on Hugging Face
0likes656downloads
Dataset Card

scene-mem-benchmark

A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work in a house for several minutes, then is asked to retrieve an object it has to remember — one that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the job it was interrupted in, remembering how far it had got — or (routine) to put a new object away where this household keeps that kind of thing, a rule it was never told and can only infer from where the robot (or someone else) put the others.

Episodes are generated in MuJoCo on AI2-THOR-derived houses with a mobile Franka Panda.

Contexts300 — type1 100 across 84 houses (find + restore) · type2 100 across 100 houses (resume) · type3 100 across 59 houses (routine; 34 twin pairs + 32 singles)
Total lengthtype1 15.4 hours (median 8.9 min, min 5.4, max 16.3) · type2 13.3 hours (median 7.9 min, min 5.0, max 13.8) · type3 14.3 hours (median 7.9 min, min 5.2, max 15.6)
Queries1772 — find 722 (388 target / 334 distractor) · restore 575 (322 restore_target / 253 restore_distractor) · resume 161 (100 resume1 / 61 multi, plus 322 ablation rows) · routine 314 (214 routine_direct — 174 from a rule the robot carried out, 40 from a rule it only saw — / 100 routine_multi)
Robot eventstype1 384 / 461 succeeded (83.3 %) · type2 every unit and errand succeeded (episodes with a failure were discarded) · type3 355 / 426 units succeeded (83.3 %) — failed units are kept; a rule counts as established once one of its units succeeded
Cameras3 × 640×480 @ 25 fps (two exocentric, one wrist)
Size114 GB contexts (type1 40 + type2 34 + type3 40) + 4 MB eval
v1.1 (2026-09-04): moved distractors. In 47 contexts one to eight distractor objects — objects the robot never touches — are moved by someone else while the robot is out of the room: seen at one place early in the episode, later seen somewhere else (some of them twice). The robot's trajectory, actions and every other object are byte-identical to v1; only the frames in which those objects (or their old spots) are visible were re-rendered from the recorded simulator state, and traj.h5 / states/final.npz carry the new poses. Such contexts have meta.json.schema_version = "1.1" and a relocated list (object, move frames, positions). The new queries are restore rows with kind = "restore_distractor" (139 first + 57 previous single rows, plus 57 multi rows that join two or three of a context's first distractors with and (one row per context, and a second two-object row in the 15 contexts with four or more) — "Pick up the A and the B and place them in or on where they first were"; in 47 of those contexts — the other moved objects failed the pick/place eligibility gates and get no query, but are moved in the video all the same); the existing rows are now kind = "restore_target" (unchanged otherwise). find is unchanged. v1.2 (2026-09-11): the `routine` family is new. 100 type3 contexts (seeds ≥ 4000) and eval/routine/ are added; nothing already on main changed. The robot is told "Put the fork away." and puts it where this household keeps cutlery — the rule (cutlery → sofa, say) is never spoken. 40 of the 100 contexts contain a rule the robot never carried out and only saw (objects already sitting on their receptacle). Contexts are generated as twins (same house, same question, rule values permuted so the answers differ); 68 of the 100 ship with their twin, 32 without. Every context ends with the robot driving back to the start table and standing in front of it (a 10–77 s tail, median 26 s), so the query starts where the holdouts are. See the routine section below. v1.1 is still reachable as revision="v1.1". v1.1 (2026-09-08): the `resume` family is new. 100 type2 contexts (seeds in the 2000s) and eval/resume/ replace the 12 provisional type2 contexts of v1 (seeds 1200/1201), which were removed from main together with their queries; the old design is only reachable at revision="v1". v1 is still reachable as revision="v1" (a git tag on this repo).
This is v1 and the layout changed. v0 shipped a flat data/<episode>/ tree with a per-episode eval_spec.json. v1 splits the release into a heavy immutable layer (contexts/) and a light layer that is regenerated whenever the query rules change (eval/). v0 is still reachable — its 153 episodes live under data/ at revision 582e691dc0f975a446d72083593b3f941a0bd3f4. Pass revision="582e691dc0f9…" (the full sha) to any huggingface_hub call: ``python from huggingface_hub import snapshot_download snapshot_download("Keh0t0/scene-mem-benchmark", repo_type="dataset", revision="582e691dc0f975a446d72083593b3f941a0bd3f4") # v0 ` The data/ path is not maintained on main` any more.

Layout

contexts/type1/<house>__<seed>/       # heavy, immutable
├── videos/{exo_camera_1,exo_camera_2,wrist_camera}.mp4
├── traj.h5                           # actions + proprioception, per simulation step
├── subtasks.json                     # two-level subtask segmentation + instructions
├── states/final.npz                  # MuJoCo state at context end — seeds the eval episode
└── meta.json

contexts/type2/<house>__<seed>/       # heavy, immutable — `resume` only (seeds ≥ 2000)
├── videos/…  traj.h5  meta.json      # same formats as type1
├── subtasks.json                     # + tasks[] (the commands), errands[], commands[] (utterance timeline), cuts[]
└── states/cut_NN.npz … final.npz     # one snapshot per unit boundary / errand end — each query starts from one cut

contexts/type3/<house>__<seed>/       # heavy, immutable — `routine` only (seeds ≥ 4000; seed 2k / 2k+1 are twins when both ship)
├── videos/…  traj.h5  meta.json      # same formats as type1; meta.json adds n_rules / n_observed / twin_seed / rules[]
├── subtasks.json                     # type2 format (tasks[] with the user's commands, cuts[]) — same runner
└── states/final.npz                  # context end — every routine query starts here (no cuts)

eval/                                 # light, regenerated when rules change (4.5 MB)
├── index.json                        # family ↔ context reverse index + row counts
├── find/{queries.jsonl, stats.json, README.md}
├── restore/{queries.jsonl, stats.json, README.md}
├── resume/{queries.jsonl, ablation_cue.jsonl, stats.json, README.md}
└── routine/{queries.jsonl, stats.json, README.md}

type1 / type2 name the context kind — the file list, not the task. A single context serves several query families: find and restore both cover all 100 type1 contexts and read the same videos, traj.h5 and states/final.npz. resume has its own 100 type2 contexts, which carry the user's utterances and per-cut state snapshots. routine has its own 100 type3 contexts, which carry the user's commands but no cuts.

Frame indices in subtasks.json and in the query rows index the videos and traj.h5 alike; everything is on the same 25 fps clock.

Download just the questions (4 MB) to see the whole benchmark before pulling any video:

python
from huggingface_hub import snapshot_download
root = snapshot_download("Keh0t0/scene-mem-benchmark", repo_type="dataset",
                         allow_patterns="eval/*")

The query families

Four families ship. find and restore are graded on the same 100 type1 contexts and both start from states/final.npz: find asks where is it now; restore asks where did it come from, and requires acting on the answer rather than just reaching it. resume (below) is graded on the 100 type2 contexts and starts from a mid-episode cut: you were interrupted — finish what you were doing, which the agent can only do if it remembers what it was doing and how far it had got. routine (below) is graded on the 100 type3 contexts and starts from the context end: put this new thing away — where, is a household rule the agent must have inferred from the episode. 322 of the 575 `restore` rows (56 %) name an object that also answers a `find` row — on those the two families are not independent measurements, and each such row says which find rows it shares an answer with in shares_answer_with. The other 253 rows (kind: "restore_distractor") are about objects the robot never touched, which moved on their own; find does not ask about those, so those rows are independent of find. On them shares_answer_with is empty for single rows and, for the 57 multi rows, lists the restore_distractor single rows whose answers the multi row re-uses.

eval/find/queries.jsonl

One line = one scoring unit. Every row carries the same 11-field envelope so a runner can read it without knowing the family, followed by the family's own fields at the same level (no nested payload).

jsonc
{"schema_version": "1.0", "family": "find", "kind": "find_target",
 "query_id": "find:train_1299__2:q00",
 "context": "contexts/type1/train_1299__2",   // dataset-root-relative
 "context_frames": [0, 17837],                // closed interval the policy may see
 "start_state": "states/final.npz",           // relative to `context`
 "query": "Pick up the cup I moved earlier",
 "budget_steps": 1500, "oracle": null, "tags": [],
 "template": "Pick",                         // wording matches how the row is graded
 "steps": "single", "name": "cup", "n_candidates": 1, "twin_nearby": false,
 "answers": [...], "candidates": [...]}

eval/<family>/README.md is the generated field dictionary; eval/<family>/stats.json is the single source for every number quoted here.

The two axes

`kind` — where the answer lives.

  • —find_target (388) — the object was moved by the robot during the context. Answering requires remembering the manipulation, not just the initial scene. This includes objects the robot moved away and later put back: they end the context at their starting place, so an agent that memorised the spoken instructions ("place it on the dresser") answers with the wrong room, and only an agent that watched the second move gets it right.
  • —find_distractor (334) — the object was never touched by the robot; it was only seen in passing. This tests incidental scene memory and guards against agents that only track what the robot handled. Objects the robot tried to pick and failed on are not distractors — no query is generated for them at all, so this kind means what it says. At most 3 distractor rows per context are kept, which is what holds the two kinds at 1:1; without the cap distractors outnumber targets 1.7:1 and are almost all plain pick.

`template` — the instruction wording, chosen to match how the row is graded. Both forms are taken verbatim from the upstream task suite this data is generated with, so the wording an agent sees at eval time is the wording it sees in training segments.

  • —Pick (673) — Pick up the {name}, from pick_task.py. Success requires lifting the object. This includes every scoring: "open_pick" row: the object is inside a closed fixture, so the fixture must be opened first, but the task still ends in a pick. The instruction does not name the fixture — saying "open the dresser and…" would hand the agent half the answer to a memory question, and its absence elsewhere would equally reveal that an object is in plain view.
  • —NavToObj (49) — Navigate to the {name}, from nav_task.py. Used only where every answer is scoring: "navigate_only": no grasp exists for that object even with fixtures opened, so reaching it is the whole task. Instructing "pick up" there would ask for something the grader does not measure and the scene does not permit.

Multi-answer rows say all the {plural} rather than upstream's any {name} ({n} available): any turns duplicate objects into a discount instead of a difficulty, and the parenthetical count leaks how many answers there are. The count stays in n_candidates.

`steps` — how many objects answer the query.

  • —single (499) — one answer.
  • —multi (223) — every answer must be picked up once within a single rollout. Order is not scored. Two flavours, told apart by tags (and by stats.json's multi_kind axis):
  • —same-name (91) — all the {plural}: several identical objects answer one name.
  • —cross-category (132, tags: ["cross_multi"]) — two or three different objects that each also answer a single row of the same context. find_target rows chain the single clauses with then, keeping each clause's own verb ("Pick up the pen I moved earlier, then navigate to the egg I moved earlier"); find_distractor rows join two or three names in one clause with and. Two- and three-object rows are split evenly (67 / 65). name joins the names with +, shares_answer_with lists the reused single rows (the measurements are not independent), and budget_steps is the sum of the singles' budgets (3000–5500).

Answer-level labels

The labels that can differ between answers live on the answer, not on the query — one query may point at an object on a counter and another inside a drawer, so folding them into a single query-level value would be lossy.

fieldmeaning
answers[].objbody id. Grading always compares body id, never category — scenes routinely contain identical twins 26 cm apart
answers[].scoringpick (822) success requires actually lifting · open_pick (78) the object sits inside a closed fixture, so it must be opened first · navigate_only (114) no grasp exists even with fixtures opened, so reaching it counts
answers[].concealmentopen (904) in plain view · hidden (110) inside a fixture that must be opened. Every hidden answer is verified to be behind a door that is still closed at context end
answers[].misplacedno (941) / wrong_surface (48) / on_floor (25) — whether the object ended up where the spoken instruction implied. An object that was moved and later put back reads no, because its last move did land where that instruction said; that it returned to its origin is visible in subtasks.json, which lists every event with its obj and settled_xyz
twin_nearbyan identical twin sits within 26 cm — a separate axis from steps
answers[].visible_at_startthe object is actually on camera in the frame the policy starts from — true on 237 of 1014 answers (160 target / 77 distractor; answer counts include the copies inside cross-category multi rows). Those rows are answerable by perception alone, so they do not test memory; decide what to do with them before reporting a number (drop them, or report them as a separate stratum)
answers[].visible_px_at_starthow many pixels it covers there, max over the three cameras, with occlusion and clipping already applied. The boolean is this >= 50, the same threshold observed/seen_s use; tighten it per row without a rebuild — >=200 keeps 175, >=500 keeps 118

hidden is the scarcest and most demanding axis: 110 answers across 55 contexts (63 outside the cross-category multi rows). Most exist because the robot successfully placed an object into a fixture and the door was still closed at the end — 90 store_in events were attempted, 51 succeeded, and every success became a query.

eval/restore/queries.jsonl — put it back where it was

575 rows across 100 contexts (5.75 per context, 1–15). The instruction names an object and a remembered place, never a described one:

Pick up the remote control and place it in or on where it first was
Pick up the spoon and place it in or on where it was just before

The target location is never spoken, so the only way to reach it is to have watched the robot move the object. Grading is a state goal: at judging time every answer must sit within r_place = 0.15 m of its own origin. Orientation is not scored (check_quat: false), and for steps: multi the order of the clauses is not scored either.

jsonc
{"schema_version": "1.0", "family": "restore", "kind": "restore_target",   // or "restore_distractor" (v1.1)
 "query_id": "restore:train_1299__2:q00",
 "context": "contexts/type1/train_1299__2",
 "context_frames": [0, 17837], "start_state": "states/final.npz",
 "query": "Pick up the remote control and place it in or on where it first was",
 "budget_steps": 5900, "oracle": null, "tags": ["store_in"],
 "template": "PickAndPlace(restore)",
 "steps": "single", "origin": "first", "name": "remote control",
 "n_candidates": 1, "twin_nearby": false,
 "shares_answer_with": ["find:train_1299__2:q01"],
 "answers": [{"obj": "remotecontrol_…_1_0_4",
              // where it is now — inside a closed fixture, so it must be opened first
              "from": {"xyz": [12.197, 12.389, 0.659], "room": 4, "concealment": "hidden", ...},
              // where it must go back to
              "to":   {"xyz": [14.679, 9.456, 0.845], "room": 4, "surface": "stand_…_1_0_4",
                       "placeable": true, "d_nearest_obj": 0.033, "d_nearest_then": 0.033,
                       "occupied_since": false, ...},
              "d_xy": 3.842, "same_room": true, "n_hops": 1,
              "r_place": 0.15, "check_quat": false, "approach": "navigate", ...}]}

from is where the object sits at context end — the same fact find grades. to is where it must go back to. Both are full location records (room, support surface, containing fixture, concealment, reach margin), so a row can be sliced on either end.

origin — which "original" is asked for

  • —first (259) — where the object stood at frame 0, before the robot ever touched it.
  • —previous (63) — where it stood just before its last move, for objects the robot moved more than once.

These form a contrast pair on the same object: an agent that only remembers the initial scene scores on first and fails previous, and one that only remembers the most recent move does the reverse. answers[].other_origins lists the object's other origins so a wrong-origin answer can be diagnosed instead of just counted wrong.

The other axes

Row-level unless marked (answers) — a multi row carries two or three answers, so the answer-level counts total 771 rather than 575.

axis
stepssingle (439) one clause · multi (136) two or three answers in one rollout, order unscored — 79 restore_target rows join same-category objects into a plural clause and different categories with then; 57 restore_distractor rows join two or three distractors with and in one clause
approach (answers)navigate (607) the destination is out of reach from where the object is picked up · in_reach (164) it is not
same_room (answers)destination in a different room (555) or the same one (216)
n_hops (answers)how many times the object was moved before the query, not rooms crossed: 1 (442) · 2 (329). A same_room answer can still read 2
from_concealment (answers)the object must first be retrieved from a closed fixture — hidden (56) · open (715)
twin_nearbyan identical twin within 26 cm of the object (171 rows) — grading compares body id, not category
d_xy (answers)how far the object must travel: median 5.9 m, p90 10.3 m, max 16.7 m

budget_steps is 5000 for a plain single-clause row, +900 when a fixture must be opened, and +3500 per extra clause on multi.

One instruction, one referent

The instruction says the {name} and nothing else, so a row only ships when that phrase picks out exactly the objects being graded. Scenes routinely hold several objects of one category, and where it first was narrows the field but does not close it — it implies only that the object moved. Three cases, resolved before a row is written:

the scene holdswhat ships
one moved object of that namea single row
several, all of them gradeableone plural clause — Pick up the mugs and place them in or on where they first were. Assignment between identical objects is not scored; the goal is that both end up home
several, some not gradeablenothing. There is no answer to fold the odd one into, so the name is dropped (16 rows)

origin: previous never merges — a plural clause cannot mix "where it first was" with "where it was just before" — so a previous row ships only when its name has a single moved bearer.

Objects that moved less than 1 m do not count against this: their origin is where they already are, so the instruction has nothing to confuse.

Eligibility

Every row is machine-verified to be doable, in this order, and rows that fail are dropped rather than shipped with a flag:

gaterequirement
pickablethe object can actually be grasped where it now sits — planned, not assumed
displacementit moved ≥ 1.0 m, so putting it back is not a no-op
origin_clearthe origin was not occupied after the fact — see below
observedthe origin was on camera ≥ 1.0 s (≥ 50 px) while the object was still there
placeablea place plan exists for that exact origin — base pose plus arm solution

172 candidate rows were dropped (116 ineligible, 34 folded into a multi row, 18 ambiguous, 4 over the 3-answer cap), a 18.9 % eligibility drop rate (the cap and the fold are not counted as drops). Per-row reasons live in the generation artifacts, which are not published; the aggregate is in stats.json.

origin_clear compares two measurements, both carried on the answer: d_nearest_then — the nearest other object at the moment this object was lifted away — and d_nearest_obj, the same distance in the end state. A row is dropped only when the origin was clear (≥ 0.20 m) and is not any more, which is the situation the generator's decoy events create deliberately: move an object, then park a different instance of the same category on the spot it vacated. A spot that was always tight is not occupied — the object demonstrably sat there — and whether it can still be placed there is what placeable measures directly. occupied_since on the answer records the verdict.

`placeable` is optimistic. The plan is built from an ideal grasp — the pose plan_pick solves for, not the pose an agent's gripper actually ends up with. It says a solution exists, not that placing will succeed. Expect the measured oracle rate to sit below the 94.2 % of candidates that passed.
Some answers are visible at t = 0. The eval episode starts from states/final.npz, which is the last context frame, and the exocentric cameras are bolted to the base — so before the policy takes a step it sees exactly what the last frame showed. On 160 of 685 `find` answers and 225 of 771 `restore` "from" positions the answer object is on camera right then (≥50 px), mostly because the robot finishes standing where it last put something down. Those rows can be solved by looking rather than remembering. Every answer carries visible_at_start and visible_px_at_start so a runner can exclude them or score them as a separate stratum; the numbers here are measured by re-rendering the start state, not inferred.
`r_place` is not yet final. 82 of 771 answers have another object closer than the 0.15 m radius (median 0.304 m over all answers), so a rollout that parks the object on that neighbour would score. Every answer carries d_nearest_obj, so a grader can tighten the radius per row without the dataset being rebuilt. This matters most on twin_nearby rows, where the twin's own home is the plausible wrong answer.

restore_distractor rows (v1.1)

Same sentence, same grading (r_place 0.15 m state goal) and the same origin = first | previous pair as restore_target; the difference is who moved the object. answers[].misplaced = "moved_by_other" and answers[].tags contains moved_distractor (plus restaged_start when the object's first place was itself chosen by us rather than the scene's original spot). answers[].moved_frames lists the frames at which the object jumped; answers[].from.settled_frame is the last of them. The old spot is empty for the rest of the episode and the new spot is seen by an exo camera for ≥ 1 s, at least 60 s and one task later (two-hop objects skip a whole task on one of the hops). Eligibility (pickable at the current spot, placeable at the origin) is a planning result as for restore_target.

eval/resume/queries.jsonl — go back and finish what you were doing

In a type2 context the user gives the robot a repetitive task ("Pick up four of the apples from the desk and put them into the basket."), then interrupts it part-way with an errand ("Hold on — go close the top drawer of the dresser in the bedroom first."), and after the errand either sends it back ("Okay, continue what you were doing.") or starts a second task while the first is still unfinished. Every transition is a user utterance — the robot never returns to an unfinished task on its own, so the video never demonstrates the behaviour being tested. The episode ends with one or both tasks unfinished.

A row is a cut: a frame at a unit boundary or right after an errand, with the MuJoCo state at that frame. The agent sees the video up to the cut (context_frames = [0, f_cut]) and is dropped into states/cut_NN.npz with one of two sentences:

`kind`pending tasks at the cut`query`rows
resume11You were interrupted earlier. Go back and finish what you were doing.100
multi2You were interrupted earlier. Go back and finish all of what you were doing.61

Neither sentence names the task, the count, or the progress; multi only leaks that there is more than one. The identity-revealing variants ("… finish what you were doing — putting the apples in the basket", and "… four in total") are shipped separately as ablation_cue.jsonl (322 rows, cue = task_named | task_named_n, role = "ablation"), one sibling row per main row with the same answer.

What the row carries — everything is the value at that cut; two rows of the same context are different moments with different answers, so never read progress from subtasks.json instead:

  • —pending[] — the unfinished tasks: n (as commanded), m (units done so far), r = n − m (the answer), done_units[], and for object tasks remaining_supply (how many are still on the source surface — always more than r) and pre_existing.dst (how many identical objects were on the destination before the robot started — counting the destination gives m + dst, not m); for visit_rooms the rooms already visited (visited_rooms) and the rooms still to visit (remaining_rooms: rid, type, and the room anchor as [x, y, z], z = floor; listed by rid — visit order is not scored).
  • —finished[] — errands (and, in variant C, the first task) already completed before the cut. Doing any of them again is a failure — a closed drawer must stay closed, a check errand must not be redone.
  • —commands[] — every user utterance before the cut (task_start, stop_errand, resume), so the context video needs no separate transcript and nothing after the cut leaks in.
  • —cut_site (unit 68 / errand_end 93), axes.locality (at_task 93 — the robot is in the room of the task it most recently worked on / away 68 — it is somewhere else, typically where the errand ended, so the unfinished job is not on screen), axes.variant (A/B/C schedule, see below), axes.interference (what happened between the task's last unit and the cut).
  • —oracle — the scripted robot's cost to finish everything from this cut: steps, unit_order, completed (source = tail: a continuation rolled out from the episode end; episode+tail: the episode's own post-cut units plus that tail). budget_steps = max(400, 1.5 × oracle.steps) (median 4410).

Grading — per pending task right_task ∧ no_redo ∧ attempt_count == r ∧ stop_correct, and for multi every pending task must pass; additionally no_finished_redo over finished[]. Order between tasks is free. attempt_count counts pick/press/visit attempts, so a slipped grasp is not a memory failure; units_completed is the secondary metric. visit_rooms is scored by the base entering the room polygon (the robot's 360° look-around on arrival is what makes a visit visible in the video, not a scoring condition). The full field dictionary is eval/resume/README.md.

Schedules. Each context has two of the four unit tasks (put objects in a basket · move objects to another surface · press an appliance button n times · visit every room), in different rooms, and 2–3 errands drawn from five primitives (check 115 · fetch_one 47 · open_fixture 44 · close_fixture 23 · close_laptop 8). Variant A (24 contexts): T1 is interrupted once and never resumed before T2 starts. B (37): T1 is resumed, then interrupted again, then T2 starts. C (39): T1 is resumed and finished, so at the final cut only T2 is pending and T1 is a decoy in finished[] (C contexts yield resume1 rows only). Pending task types over all rows: putincontainer 67 · visitrooms 62 · pressbutton 49 · move_objects 44; total remaining units per row 1–5 (median 2).

Progress is hidden by construction. The source surface holds n + k_src objects and the destination starts with k_dst identical ones, so neither counting what is left nor counting what has arrived gives m; button presses and room visits leave no trace at all (each press is reach → touch the panel → retract → fold the arm back to its home pose, so every press looks the same). Rows where the remaining supply did not exceed r, where a cut was not restorable to < 1 mm, or where the oracle could not finish were dropped, and contexts in which a second object was dragged along with the one being carried, or in which an errand had nothing left to do, were discarded in generation.

eval/routine/queries.jsonl — put it where this household keeps it

In a type3 context the user asks the robot, four times per rule, to put something away ("Put the fork away."), and the robot carries it to the household's place for that kind of object — a rule of the form object class → receptacle class (cutlery → sofa, food → desk, gadgets → basket). The rule is never spoken: it is only visible in where things end up. Values are drawn at random, so common sense is right in at most one of the K receptacles. Each context has 2 or 3 rules over six object classes (FOOD, DISH, CUTLERY, GADGET, CLOTH, TOILETRY) and eight receptacle classes (basket, sofa, desk, dining table, side table, counter top, coffee table, shelving unit). In 40 contexts one rule is observed, not executed: three objects of that class already sit on their receptacle and the robot only passes them on its way (rules[].evidence_source = "observed", with the frames and pixel counts in which they were on camera). A holdout — a sub-category of each class that never appears in the episode (a hand towel when the robot moved cloths; a tomato when it moved food) — sits on the start table from the first frame and is never touched.

After the last unit the robot returns to the start table and stands facing it (subtasks.json return_to_table / tail give the frames; meta.json.tail repeats them), so the context ends where the holdouts are. A row is a question about the holdouts, asked at that moment (start_state = states/final.npz, context_frames = the whole video):

`kind``query`answerrows
routine_directPut the hand towel on the countertop away.that object on any instance of its rule's receptacle class214 (174 executed rule · 40 observed rule)
routine_multiClear the countertop.every holdout on the table (plus whatever else the robot left there) on its class's receptacle100

Neither sentence names a receptacle. answers[] carries receptacle_class, any_instance_ok = true and instances[] (the scene bodies of that class); grading is by class, scoring = "result": the object must end up on (support_of) or in (basket) one of those instances. Rows are counted per rule that was actually established (≥ 1 successful unit for an executed rule; ≥ 1 unit with all three objects on camera for an observed one), and contexts whose end state contradicts a rule (an object of a class resting on another rule's receptacle) were discarded.

Twins. Contexts are generated in pairs (seed 2k and 2k+1 in the same house, twin_context): same house, same holdouts, same questions, rule values permuted without a fixed point — every answer differs, so a policy that guesses from common sense (or copies the twin) is wrong in one of them. 68 of the 100 contexts ship together with their twin and 216 of the 314 rows have a twin_id (the paired row); grade those as pairs. The other 32 contexts (twin_id = null on all their rows) are shipped alone because their twin failed generation — this cut prioritises the 40 observed-rule contexts (the family's hardest and rarest kind: only 11 of the 63 generated ones have a surviving twin) over pair completeness.

Fields on every row: rules[] (the established rules with evidence frames), holdout[], answers[], twin_id, twin_context, axes (n_rules, n_observed, evidence_source, holdout_subcat; n_items, n_unseen on multi), counterexample (always null on shipped rows), cue = "bare", rule_stmt = null. budget_steps and oracle are null (see Known issues). The field dictionary is eval/routine/README.md.

Selection

type1 contexts were filtered by hard gates, applied in order, and nothing was selected by hand:

gaterequirement
filesall of traj.h5, subtasks.json, meta.json, the query spec, the state snapshot
videosall three cameras present and non-empty
statea real MuJoCo state (qpos/qvel/ctrl/time) — an eval episode cannot start without it
scoredevery query passed eligibility (no scoring: "unknown" left)
failed≤ 1 failed robot event — with 4–5 events per episode that means a 75–80 % success rate
length≥ 5 minutes
queries≥ 5 scorable queries

The failed ≤ 1 gate is by far the strictest: it is what takes the pool from 548 to 156. It also has a cost worth knowing before you use the misplaced axis — objects end up in the wrong place because the robot failed, so a clean-episode filter thins that axis to 54 contexts. If your work needs failure cases, the generation side can produce a looser cut.

Queries that failed validation were removed rather than shipped with a flag, so every row in `queries.jsonl` is scorable. 228 were dropped: 150 capped (at most 3 find_distractor rows per context — distractors otherwise outnumber targets and are almost all plain pick, so the cap keeps the two kinds close to balanced and costs no rare axis), 57 ambiguous (a distractor query's wording collided with a target's), and 21 ineligible — mostly because the robot never actually observed the answer object (asking "the one from earlier" about something that never appeared on camera is not a memory question), plus 3 objects the robot tried to pick and failed on, leaving them exactly where they started: neither kind is honest about those, so no query is made. Per-query reasons live in the generation artifacts, which are not published; the aggregate is in stats.json.

type2 contexts went through their own gates: every unit and errand succeeded (a failed unit, a failed errand, an object dragged along with the carried one, or an errand that turned out to have nothing to do discards the episode), real length ≥ 300 s, every cut snapshot restores to < 1 mm, and every row's oracle finishes. Of roughly 575 rollouts 180 completed and 135 passed the gates; the 100 shipped were chosen to balance the task pairs and schedule variants (the surplus is kept aside, not published).

type3 contexts: real length ≥ 300 s, every established rule has ≥ 1 successful unit (executed) or ≥ 1 unit with all three evidence objects ≥ 200 px on camera (observed), the holdouts are still on the start table, no counter-example in the end state (checked by rebuilding the scene from states/final.npz), and no object dragged along with the carried one. Failed units are kept (71 of 426): a household rule is still visible when 3 of 4 units succeeded. Of 304 passing episodes the 100 shipped are: the 40 observed-rule episodes with the most rows (those with a surviving twin first), then 30 complete twin pairs without an observed rule, by row count. The surplus (204 episodes, among them 112 complete pairs) is kept aside, not published. Every context ends with the robot returning to the start table: that tail was simulated and rendered separately from the context's final state and appended (traj.h5, the three videos and states/final.npz continue seamlessly; subtasks.json.tail marks the seam).

Leakage — contexts/ always holds the whole episode

Every file under contexts/ covers the episode end to end. For find and restore that is harmless (context_frames spans the full episode in both), but the field is part of the contract and a runner must honour it: the reader truncates, not the dataset. For resume it is essential: the frames after the cut show the next command and the remaining units — the answer — so the video and traj.h5 must be cut at context_frames[1], and subtasks.json must not be consulted for progress (the row's pending/finished/commands carry everything the agent may know). Families that cut mid-episode reuse the same context files rather than shipping a clipped copy.

traj.h5

One group per trajectory (traj_0), ManiSkill trajectory format:

traj_0/
├── actions/{commanded_action, ee_pose, ee_twist, joint_pos, joint_pos_rel}
├── obs/agent/{qpos, qvel}
├── obs/extra/{env_states, policy_phase, policy_num_retries, …}
├── obs/sensor_param/<camera>/{intrinsic_cv, extrinsic_cv, cam2world_gl}
├── env_states/articulations/panda
└── {rewards, success, terminated, truncated, fail}

Camera intrinsics and extrinsics are stored per frame, so the exocentric views can be lifted to world coordinates and lined up against ground-truth object poses. Images are not in the h5 — they are in the mp4s, at the same frame indices.

subtasks.json

Two levels, with deliberately different field names so training targets do not blur:

  • —L1 `subtasks[]` — one entry per event, using task_description. Navigation to the docking pose is inside the subtask, so the count matches the number of tasks in the context.
  • —L2 `segments[]` — cut at base motion, using instruction taken verbatim from the upstream task templates. Three kinds only: nav (moving empty-handed), carry (moving while holding), manip (arm only, base parked). manip segments carry robot_base_pose.

outcome on each subtask records what actually happened (ok, failed_dropped, failed_no_plan, …) — failures are kept, not hidden.

type2 files use schema 2.0 of the same layout: L1 rows are the task units (t1u0, …) and the errands (e0, …), plus tasks[] (the commands, with utterance, start_frame, n_planned, n_done), errands[], commands[] (the utterance timeline with frames) and cuts[] (every snapshot with its progress). Read these for analysis only — for grading, the query row is the contract.

Loading

python
from huggingface_hub import snapshot_download
import json, h5py

root = snapshot_download("Keh0t0/scene-mem-benchmark", repo_type="dataset")

rows = [json.loads(l) for l in open(f"{root}/eval/find/queries.jsonl")]
rows += [json.loads(l) for l in open(f"{root}/eval/restore/queries.jsonl")]
for r in rows[:5]:
    print(r["family"], r["query"], "->", [a["obj"] for a in r["answers"]])

resume = [json.loads(l) for l in open(f"{root}/eval/resume/queries.jsonl")]
r = resume[0]                      # start state = states/<cut>.npz, video up to context_frames[1]
print(r["query"], r["context_frames"], [(p["task_id"], p["type"], p["r"]) for p in r["pending"]])

routine = [json.loads(l) for l in open(f"{root}/eval/routine/queries.jsonl")]
r = routine[0]                     # start state = states/final.npz; twin_id names the paired row
print(r["query"], "->", [(a["obj"], a["receptacle_class"]) for a in r["answers"]], r["twin_id"])

r = rows[0]
with h5py.File(f"{root}/{r['context']}/traj.h5") as f:
    print(f["traj_0/obs/agent/qpos"].shape)

Rows are sorted by (context, query_id) within each family, so grouping by context gives you every question about one episode in one place — load the video once per group, not once per row. Group across families too: the runner's cache key is (context, context_frames[1]); find and restore produce the same value for a shared context, while the two resume rows of one context have different context_frames[1] (the earlier cut is a prefix of the later one).

To pull a single context instead of the whole tree:

python
snapshot_download("Keh0t0/scene-mem-benchmark", repo_type="dataset",
                  allow_patterns="contexts/type1/train_1299__2/*")     # or contexts/type2/train_123__2000/*, contexts/type3/train_3118__580[89]/* (a twin pair)

Known issues

  • —oracle is null on every find and restore row — the oracle policy has not been run against this cut, so there is no measured step budget and no measured ceiling yet. budget_steps is an estimate: find 1500 open / 2500 behind a door, restore 5000 + 900 per door + 3500 per extra clause. resume rows do carry a measured oracle (the scripted robot finishing from the cut).
  • —routine rows also have oracle = null and budget_steps = null — no oracle has been run from the context end; the plan's budget is max(400, 1.5 × oracle). Grading (routine_judge.judge, class-level support_of) has been used only as the generation-side counter-example check, not by a policy grader.
  • —routine: 32 contexts ship without their twin (twin_id = null on 98 rows), so the pair metric covers 216 of 314 rows; among the 40 observed-rule rows only 8 (4 pairs) are pair-gradable.
  • —resume grading has not yet been exercised by a grader; the per-task rule above is the specification (eval/resume/README.md). attempt_count needs the runner to log pick / press / room-entry attempts.
  • —restore eligibility is a planning result, not an execution one (see the note above), so the achievable success rate is below the 94.2 % of candidates that passed.
  • —restore grading radius r_place is fixed at 0.15 m and has not been validated against a grader, because no grader has been written yet. On 82 of 771 answers a neighbouring object falls inside that radius; d_nearest_obj is on every answer so the radius can be tightened per row.
  • —Whether the robot drives through walls has not been verified for this cut. The checker exists on the generation side but was not run per episode.