Team Ai
Datasetpublic

axisrobotics/Test-Dataset

AXIS held-out 20 — a cross-embodiment few-shot adaptation benchmark 20 tasks never seen in pretraining, 20 demonstrations each, LoRA adaptation, rollout evaluation. This release is the SPECIFICATION and the INDEX, not the demonstration data. It is published first and on purpose: everything here is what you need to render the benchmark on your own embodiment, and none of it depends on our video encoding being finished. Status: the task list is CANDIDATES. The learnability gate… See the full description on the dataset page: https://huggingface.co/datasets/axisrobotics/Test-Dataset.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes90downloads
README.md147 linesDownload Raw Back to render_request
1# Render request — held-out benchmark, 20 tasks × 20 demos2 3Everything needed is in this folder. Read the attempt lists here, upload the rendered output here.4 5**What:** re-render 400 existing demonstrations (20 tasks × 20 attempts) for the **default Franka**,6at the **stage1_5k camera rig**, with **appearance pinned**.7 8### Why 20 and not 10 — and why the existing LIBERO-rig renders don't count9 10Ten of these tasks are currently in the stage1_5k pretraining data, so they cannot be scored against11today's models. They are listed as `status: conditional`. The other ten are `status: active`.12 13It would seem to follow that only the ten active tasks need rendering. It does not, for two reasons:14 151. **The conditional ten become usable the moment a stage-1 run excludes the held-out set** — which16   is planned. Rendering them now avoids a second round trip.172. **Their existing renders don't meet the appearance requirement.** Tasks 810/815/929/... already18   have renders at the correct camera rig in `Franka-Datasets-v2-5k-LIBERO`. We checked those19   manifests: they run `TextureModder` and `LightingModder` too. So they need re-rendering for20   appearance regardless of the rig.21 22So the ask is all 20. If you have to stage it, do the ten marked `active` first.23 24**Why re-render:** these demonstrations currently exist only at a superseded camera rig. The25pretraining corpus (`stage1_5k_camera_fixed`) uses a different one, and this benchmark evaluates26models pretrained on it — so a mismatch would give every evaluated model a train/eval viewpoint27gap rather than measuring the model.28 29---30 31## 1. The camera rig — must match exactly32 33Verified identical across sampled tasks of `Franka-Datasets-v2-5k-LIBERO` and34`Franka-Datasets-v2-30k-LIBERO`. Machine-readable copy: [`camera_rig.json`](camera_rig.json).35 36| | value |37|---|---|38| embodiment | **franka** (default) |39| camera group | **`camera_fixed`** |40| wrist camera | mounted on the **end-effector body**, hand→camera translation **`[0.05, 0.0, 0.0]`**, quat **`[0, 0.707108, 0.707108, 0]`**, **fovy 75°** |41| third-person camera | `camera_to_world` translation **`[1.0086, 0.0, 1.1904]`** |42| resolution | **640 × 360** |43| fps | **15** |44 45**The wrist rig is the one that matters most.** The superseded batch had it at46`[-0.074, 0, 0.0292]` with fovy 51.9° — 12.7 cm away, on the opposite side of the hand, and 23°47narrower. If a render comes back at the old rig it is unusable, and the difference is invisible in48a single still frame: it only shows up in how the image moves.49 50A quick self-check after rendering one episode: compose the inverse end-effector pose with the51recorded `camera_to_world` and confirm the translation is `[0.05, 0, 0]` to 3 decimals, and that it52is constant across frames (it should be rigid — ours measures a per-element std of ~7e-08).53 54---55 56## 1b. Appearance must be PINNED — this is new, and it matters more than the rig57 58The existing demos are **appearance-randomized per demo**: the render pipeline runs59`TextureModder` and `LightingModder` off a per-attempt `variant_seed`, so each of a task's 2060demonstrations has a different table material, a different floor and different lighting. We only61found this by looking at the frames — it is recorded in the render manifest, not in the task's62`domain_randomization` block, which contains position deltas only.63 64For a 20-demo benchmark that is a problem. The policy has to learn appearance-invariance *and* the65manipulation skill from 20 samples, so the score conflates two abilities and cannot be attributed66to the one under test. Measured on the current demos: a reach task scores 82% while two67manipulation tasks score 2% and 0%, with 99 of 100 rollouts running to the step cap — acting, never68completing.69 70**So please render all 20 demos of a task under ONE fixed appearance.** Everything in this list71held constant across every demo of every task:72 73| pin | why |74|---|---|75| **table material** | varies per demo today — six demos of task 953 have six different tables |76| **floor / background** | varies with it; one demo sits on brown parquet, another on green grass |77| **lighting** | `LightingModder` is active per demo |78| **object materials / textures** | `TextureModder` recolours the OBJECTS too, so the same mesh reads as a different object between demos |79 80Concretely: disable `TextureModder` and `LightingModder`, or hold `variant_seed` constant across81the whole batch.82 83**Object meshes do NOT need changing** — we verified those are already consistent. The eval scene84loads `cup_2`, `muffin_4`, `sweet_potato_4`, `tissue_box_2` from exactly the same85`tabletopgen/task_88200032_layout_1_rand_00000` directory the training scene references. The86problem is purely that their *materials* are randomized, so identical geometry looks like a87different object from demo to demo.88 89Object PLACEMENT should keep varying — that is the axis the benchmark is meant to test.90 91If your pipeline cannot pin any of these, please say so rather than varying them — we would rather92know and design around it than discover it in the frames again.93 94## 2. What to render95 96| file | contents |97|---|---|98| [`tasks.json`](tasks.json) | all 10 tasks, required + spare attempt ids, the rig |99| [`attempts/task_<id>.json`](attempts) | one file per task |100| [`attempts/all_attempts.csv`](attempts/all_attempts.csv) | flat table — `task_id, task_name, operation, role, attempt_id, n_frames, reward_final` |101 102See [`tasks.json`](tasks.json) and [`attempts/all_attempts.csv`](attempts/all_attempts.csv)103for all 20 tasks, each marked `active` or `conditional`.104 105 106**`attempt_id` is the identity — please render exactly those.** They are not arbitrary: they were107selected as the highest-reward demonstrations available for each task, and they are the only ones108with verified trace and goal data. Episode indices are renumbered by every build, so please key109your output on `attempt_id`, not on episode order.110 111**`role = spare`** lists 40 further attempts per task, ranked. If a required attempt fails to112render, take the highest-ranked spare and say which you substituted — don't silently drop one, as11320 demos is the whole adaptation budget and a missing demo is 5% of it.114 115---116 117## 3. Where to upload118 119Please upload into `uploads/` in this same folder, one directory per task:120 121```122uploads/123  task_809/124    meta/          LeRobot v3.0 metadata125    data/          parquet126    videos/        observation.images.third_person/ and .wrist/127    render_manifest.json     attempt_id -> episode_index, per episode128  task_811/129  ...130```131 132`render_manifest.json` is the important one — without an `attempt_id → episode_index` mapping we133have to recover the pairing by matching trajectories, which works but is slower and needs checking.134If your pipeline already emits `task_manifest.json` in the usual shape, that is perfect; please just135make sure `records[]` is populated (two tasks in the previous batch shipped with `records: []` while136having over a thousand episodes).137 138---139 140## 4. Anything else141 142If a task can't be rendered at this rig for a structural reason, please say which and why rather143than substituting a different rig — we'd rather run a 9-task benchmark than one with a mixed rig.144 145Questions on any of this are welcome before rendering starts; a wrong rig is expensive to discover146afterwards.147