Team Ai
Datasetpublic

3d-layout/blender-3d-layout-v2-720p

Blender 3D Layout — 720p comparison groups (release v3.1) Synthetic multi-view renders of objects placed into real Blender scenes, built so that a change to the layout can be isolated: every camera pose is rendered for the base layout, with one object moved, with an extra object added, and with one object rotated in place — with the same camera matrix each time. Each layout also comes with a box-orientation prompt in the style of BoxCtrl: every object's 3D box drawn on a white… See the full description on the dataset page: https://huggingface.co/datasets/3d-layout/blender-3d-layout-v2-720p.

sourceHugging Faceotherupdated 1d agoView on Hugging Face
0likes4kdownloads
Dataset Card

Blender 3D Layout — 720p comparison groups (release v3.1)

Synthetic multi-view renders of objects placed into real Blender scenes, built so that a change to the layout can be isolated: every camera pose is rendered for the base layout, with one object moved, with an extra object added, and with one object rotated in place — with the same camera matrix each time. Each layout also comes with a box-orientation prompt in the style of BoxCtrl: every object's 3D box drawn on a white background with three semi-transparent coloured faces, so position, size and orientation can be read off directly.

Release v3.1 (2026-10-02) re-renders every box-orientation image on a pure white background — the faded source image under the boxes is gone, the boxes themselves are unchanged (see Changelog). Release v3.0 regenerated 206 of the 257 groups in which placed objects intersected the scene and added per-group intersection flags. The repository history was squashed on 2026-10-02: only v3.1 is on the Hub; the earlier releases (v2.3, v3.0) and their files are gone. Pin revision="v3.1" for reproducible work.

Separate subset [`rot_single_v1/`](rot_single_v1/README.md) (added 2026-10-04): the simplest rotation setting — one object alone on a pure white background, before and after a turn about the vertical axis (300 objects × 20 pairs = 6,000 pairs, same 1280×720, box-orientation images with opaque coloured faces). It is independent of the comparison groups described below and has its own README; pin the tag rot-single-v1 (= v3.1 plus this folder) or download only it with --include "rot_single_v1/*".

The layout structure is "v2": it replaces the v1 layout of 5 layouts + 5 independent deletion layouts, where the two directories did not share viewpoints (view_00 in one was not view_00 in the other). Here the variants of a group are pixel-aligned by construction.

Contents

Scenes165
Comparison groups501 (490 with a rotate variant)
Camera poses (views)6,954
Images68,946 PNG, 1280×720 (27,618 RGB, 27,618 BBOX, 13,710 box-orientation)
Object masks193,540 PNG
Distinct object assets placed211 (from a 1,000-asset pool)
Per-object text captions1,000
Groups free of intersection flags449 of 501 (metadata/clean_groups.txt)

Every view yields {base, move, add} × {rgb, bbox} plus base/orient, and — in the 490 groups with a rotate variant, for the views where the rotation is visible — rotate/{rgb, bbox, orient}. Most groups have 16 views; 164 have fewer (views where the layout was not visible were dropped), so always read the view list from scene_result.json instead of assuming 16.

Download

Use huggingface_hub 1.4 or newer (tested with 1.29; older versions treat a repeated --include differently and retry rate limits less):

bash
pip install -U "huggingface_hub>=1.4,<2"

The repo is public: no login is needed (hf auth login or HF_TOKEN only raise your rate limits).

Everything, pinned to this release, into a new, empty directory:

bash
hf download 3d-layout/blender-3d-layout-v2-720p --repo-type dataset --revision v3.1 --local-dir blender-3d-layout-v3.1
python
from huggingface_hub import snapshot_download
snapshot_download("3d-layout/blender-3d-layout-v2-720p", repo_type="dataset", revision="v3.1",
                  local_dir="blender-3d-layout-v3.1")

Subsets — glob patterns, * also matches /. Keep *.json (or at least */scene_result.json) in every subset: the images are useless without the boxes, cameras, rotation angles and the orientation legend stored there.

python
REPO = "3d-layout/blender-3d-layout-v2-720p"
kw = dict(repo_type="dataset", revision="v3.1")
snapshot_download(REPO, local_dir="sub", allow_patterns=["0001_*", "metadata/*"], **kw)        # one scene + metadata
snapshot_download(REPO, local_dir="rot", allow_patterns=["*/rotate/*", "*.json"], **kw)        # rotate variant + JSON
snapshot_download(REPO, local_dir="orient", allow_patterns=["*/orient/*", "*.json"], **kw)     # box-orientation images + JSON
bash
# the same with the CLI (1.4+: repeat --include once per pattern)
hf download 3d-layout/blender-3d-layout-v2-720p --repo-type dataset --revision v3.1 --local-dir rot \
    --include "*/rotate/*" --include "*.json"

Both list the whole repo (263,668 files) before the first download, which takes a few minutes even for a tiny subset; download_update.py download --scenes ... below lists only the matching scene folders.

Already have an older copy? A v3.0 copy is brought to v3.1 by download_update.py update below, which fetches only the 13,880 changed files and removes the two v3.0 update notes; hf download --revision v3.1 into the same directory works too (then download_update.py update --from v3.0 --prune-only removes those notes). An older copy (v2.3) can no longer be updated step by step — v3.0 is gone from the Hub — and hf download / snapshot_download into an existing directory do not delete anything (v3.0 removed 13,822 files of views that no longer exist). Run download_update.py download into it (fetches only files that are missing or differ) followed by download_update.py verify --prune (deletes local files v3.1 does not have), or download into a fresh directory.

scripts/download_update.py

Standard huggingface_hub, nothing else. It checks the sha256 of every file, resumes where it stopped, and can update a v3.0 copy in place from the manifest metadata/updates/v3.1.json — fetching only what changed and deleting what was removed (only if it still matches its old hash, and only after every download succeeded); download followed by verify --prune brings any other copy to v3.1:

bash
hf download 3d-layout/blender-3d-layout-v2-720p scripts/download_update.py --repo-type dataset --revision v3.1 --local-dir .
python scripts/download_update.py update   --local-dir ./blender-3d-layout --dry-run  # the plan only
python scripts/download_update.py update   --local-dir ./blender-3d-layout --verify   # v3.0 -> v3.1, then check every file
python scripts/download_update.py download --local-dir ./old-copy                     # any older copy: fetch what differs,
python scripts/download_update.py verify   --local-dir ./old-copy --prune             # then delete what v3.1 does not have
python scripts/download_update.py download --local-dir ./sub --scenes 0001,0084      # scene subset
python scripts/download_update.py download --local-dir ./rot --variants rotate --kinds rgb,orient
python scripts/download_update.py download --local-dir ./clean --clean-only group    # intersection-free groups
python scripts/download_update.py snapshot --local-dir ./full --revision v3.1      # everything + verify
python scripts/download_update.py verify   --local-dir ./full --revision v3.1 --prune
  • —Subsets come with their annotations: scene_result.json, masks.json, scene_config.json and group_NN/cameras.json are included with any --kinds selection (--no-json leaves them out).
  • —Updating a full v3.0 copy downloads about 608 MB (1 new and 13,879 changed files) and deletes 2 files (metadata/updates/v3.0.json / .md); the other 249,786 files are left as they are. A partial copy is updated with --existing-only (only files you have) or the same --scenes/--variants/--kinds filters you downloaded with. update reads which release the directory is from the record the script leaves in it (.dataset_revision.json) or from its README.md (--from v3.0 says it explicitly) and checks that against the download records hf/the script leave; if it cannot tell, it says how to bring the copy to v3.1 with download + verify --prune. A copy inside the Hugging Face cache (downloaded without --local-dir) is not updated in place: run snapshot_download(..., revision="v3.1"), which reuses the unchanged files.
  • —Time: the Hub allows about 5,000 file requests per 5 minutes per account, so expect at least ~1 minute per 1,000 files whatever the bandwidth — ~14 minutes for the 13,880 files of a full update from v3.0. Ctrl+C stops at once; re-run the same command to resume (files already present with the right hash are not fetched again).
  • —`verify` fails on missing files, hash mismatches and local files the revision does not have (e.g. removed views); --prune deletes the latter.
  • —Windows: the longest path in the repo is 128 characters. Keep --local-dir short (e.g. D:\bl3d) or enable long paths; prefer --local-dir over the HF cache, whose own path prefix is long.
  • —A .cache/huggingface/ folder inside the download directory holds huggingface_hub's per-file download records; keep it to make resuming fast (without it every file is re-hashed once).

Layout

<rank>_<scene_asset_id>_<slug>/
├── scene_config.json        render settings for this scene
├── scene_result.json        full metadata: groups, objects, cameras, poses
├── masks.json               per-object mask records and checks
├── worker.log               render log (provenance)
└── group_NN/
    ├── cameras.json         the shared camera set for this group
    ├── base/{rgb,bbox,orient}/view_NN.png
    ├── move/{rgb,bbox}/view_NN.png
    ├── add/{rgb,bbox}/view_NN.png
    ├── rotate/{rgb,bbox,orient}/view_NN.png
    └── */masks/…                per-object masks, see below
metadata/
├── index.json               scene → group → view index
├── object_pool.json         the 1,000-asset pool (shared, not duplicated per scene)
├── object_captions.json     per-object captions keyed by asset_base_id
├── object_captions.jsonl    same, one JSON object per line
├── captions_report.json     captioning run report
├── rotate_plan.json         the fixed 180° / random-angle assignment per group
├── rotate_qc.txt            whole-dataset QC report of the rotate variant
├── intersect_flags.json     per group × variant intersection flags (also .csv)
├── clean_groups.txt         groups with no intersection flag in any variant
├── regen_summary.json       which groups were regenerated in v3.0, which could not be
├── dataset_summary.json     counts and distributions
└── updates/v3.1.json        every path v3.1 added / modified / deleted since v3.0 (sha256, size; + .md)
scripts/download_update.py   download / subset / update / verify
rot_single_v1/               separate single-object rotation subset (own README, pairs.jsonl)

metadata/mask_summary.jsonl, mask_audit.json and rotate_*_summary.jsonl are run logs of the original v2 mask and rotate passes; they are not rewritten for regenerated groups. The per-scene masks.json totals are authoritative.

The comparison-group guarantee

Within one group_NN, base, move, add and rotate share identical camera matrices and lens for the same view_index — checked element-wise across all 6,954 group views (6,756 for rotate) of this release, 0 exceptions. So group_00/base/rgb/view_07.png, group_00/move/rgb/view_07.png and group_00/add/rgb/view_07.png differ only by the layout change.

  • —base — N objects placed on a support surface (N ∈ {2,3,4,5})
  • —move — exactly one object displaced; all others frozen
  • —add — base plus one additional object; base objects frozen
  • —rotate — exactly one object turned about the vertical axis; all others frozen (see Rotation variant below)

Because add contains a superset of base, the pair also serves as a deletion example read in the other direction.

Camera correspondence holds within a group, not across groups: each group gets its own camera set, fitted to its own object cluster.

BBOX images

bbox/ frames are the 3D oriented bounding boxes rendered as translucent coloured volumes (alpha-blended, not opaque fills) over a white backdrop, one colour per object. They are the exact projection of the bbox_obb_world.corners stored in scene_result.json — reprojecting those corners with the recorded camera lands within ~2–3 px of the rendered box.

Object masks

Every view carries per-object masks, rendered from the same scene state as the RGB frame (same asset placement, same camera) with Blender's Workbench engine — flat shading, per-object colour, anti-aliasing and dither off — so each mask is an exact rasterisation of the object, not a segmentation-model estimate.

group_NN/{base,move,add,rotate}/masks/
├── amodal/view_NN_objK.png    1-bit; object K's full projection, occluders ignored
├── bbox3d/view_NN_objK.png    1-bit; solid projection of object K's 3D bounding box
└── visible/view_NN.png        8-bit label map; 0 = background, K+1 = object K

amodal is the object alone with everything else hidden, so masks of two objects overlap where one sits behind the other and the intersection is simply amodal_i & amodal_j. bbox3d is the same idea applied to the object's oriented 3D bounding box (bbox_obb_world), rendered as a solid cuboid through the same camera — the box-shaped, amodal region a SeeThrough3D-style token mask uses; bbox3d_i & bbox3d_j gives the overlap. visible is what the RGB frame actually shows, with both scene geometry and other objects occluding.

amodal_K ⊆ bbox3d_K holds for every object and view except one asset: the LED Work Light (55bbf9ad-…) has a power cable modelled as a curve, which the mesh-vertex OBB does not enclose, so a thin run of its amodal mask lies outside its box. Both masks are rasterised through the same camera, so the check is exact rather than a reprojection estimate.

K indexes variants.<variant>.objects[] in scene_result.json; the same order is recorded in each scene's masks.json together with per-view pixel counts, pairwise amodal overlaps, and the check results below.

Every mask was checked, not sampled:

  • —visible_K ⊆ amodal_K for every object and view (same rasteriser, so a violation would be a decode error) — 0 violations outside the scenes listed under Known issues.
  • —amodal_K lies inside the projected 3D bounding box of object K (+6 px) — verifies the object was rebuilt where the RGB frame had it.
  • —For move, the pixels that differ between base and move RGB fall inside the moved object's amodal mask; what falls outside is the object's shadow and reflection moving with it (masks.json → change_check).

Masks are stored at full 1280×720. Downsample to a token grid (e.g. 16 px patches → 80×45) downstream; nothing is pre-quantised.

Rotation variant and box-orientation images

Every group also carries a rotate variant: one base object turned about the world vertical axis through the centre of its oriented box, everything else frozen, rendered through the group's same cameras. It isolates a change of orientation the way move isolates a change of position.

group_NN/
├── base/orient/view_NN.png              box-orientation image of the base layout
└── rotate/
    ├── rgb/view_NN.png                  the rotated layout (prediction target)
    ├── bbox/view_NN.png                 same BBOX style as the other variants
    ├── orient/view_NN.png               box-orientation image of the rotated layout
    └── masks/{amodal,visible,bbox3d}/   as for the other variants

Box-orientation image (BoxCtrl-style prompt). Modelled on the RGB 3D box prompts of BoxCtrl (Wang et al., SIGGRAPH 2026): every object's bbox_obb_world is drawn through the view's render camera on a pure white background (255,255,255) — the source image is not drawn (in v3.0 it was, faded to 35%). The rotate prompt draws the rotated layout's boxes the same way. Three faces of the box's own local frame are painted:

facelocal axiscolour (sRGB), 50% opacity
front−Yred (220,40,40)
top+Zblue (45,70,230)
left−Xgreen (40,170,60)
back, right, bottom+Y, +X, −Zfully transparent

The coloured faces are drawn whichever way they face, so a face turned away from the camera still shows through the box — after a rotation you can see where the red face went. All 12 edges are black (~2 px). Faces and edges are painted far to near, so an edge behind a coloured face is tinted by it. Face names follow Blender's views (Front looks at the −Y face). The local frame is the asset's authored frame: for the assets checked by eye (screens, monitors, arcade cabinets) −Y is the visual front, but it is the authored axis, not a per-asset semantic label. The exact style is recorded in scene_result.json → orientation_image.

The prompts are rasterised from the stored boxes with the scene's actual render camera (render_camera, below). Checked against Blender's own renders of the same boxes on the 14,090 prompts of the pre-regeneration set: median silhouette IoU 0.963 (the gap is edge thickness), including orthographic and panoramic cameras; the prompts of regenerated groups come from the same renderer.

Angles. 147 of the 490 rotated groups (30.0%) are turned by exactly 180° — the 30% split was drawn once over all groups with a fixed seed before rendering. The rest are turned by |θ| uniform in [30°, 150°] with a random sign (positive = counter-clockwise seen from above, Z-up).

Which object. Drawn at random among base objects that are clearly visible (≥ 1,500 visible pixels and ≥ 30% of the amodal mask, in at least half the views); if none qualifies, the most visible objects are tried in order. A non-180° turn is rejected if the rotated footprint would overlap another object's footprint, lose support on its surface, or intersect scene geometry; a 180° turn about the box centre maps the box onto itself, so it needs no such check.

What is guaranteed. rotate differs from base only by the rotated object:

  • —Before rendering a group, its first base view was re-rendered from the rebuilt base state and compared with the stored frame (8/255 tolerance): median 0 changed pixels, 99th percentile 2,869, maximum 15,245. Groups whose base did not reproduce were withdrawn (rotation.status = "failed").
  • —A view is kept only if at least 200 pixels change inside the rotated object's before/after box projection; 77 views where the turn is hidden or out of frame were dropped, so variants.rotate.views can be a subset of the base views (rotation.dead_views lists them).
  • —Changed pixels far outside the rotated object's box beyond the scene's own render noise would mean a frame rendered in a different scene state: 4 views (1 groups), all physical — e.g. a rotated table lamp relighting the table next to it, a guitar's reflections on a glossy floor.
  • —Rotate masks pass the same checks as the other variants: 0 containment violations over 18,596 masks; the only amodal ⊄ bbox3d cases are the LED Work Light cable described above.

Metadata (scene_result.json):

  • —orientation_image — the colour legend above.
  • —groups[].rotation — object_index, asset_base_id, object_name, angle_deg, is_180, axis, pivot_world, base_consistency, dead_views, views_kept, and tried (candidates rejected, with reason).
  • —groups[].variants.rotate.objects — the base objects, with the rotated one's empty_matrix_world, bbox_obb_world (corners, frame, centre), bbox_aabb_world, yaw and yaw_delta_to_base (radians) updated.
  • —groups[].variants.rotate.views[] — rgb_path, bbox_path, orient_path, the shared camera, and rendered_change_px / rendered_change_roi_px (changed pixels in the frame / inside the object's box).
  • —groups[].variants.base.views[].orient_path.

In masks.json the rotate masks sit under groups[].variants.rotate like the others, with their totals kept apart in rotate_totals and the base→rotate change check in groups[].rotate_change_check.

Camera parameters

Each camera record carries full intrinsics, not just focal length:

json
"intrinsics": {
  "type": "PERSP", "width": 1280, "height": 720,
  "fx": ..., "fy": ..., "cx": ..., "cy": ..., "K": [[...]],
  "lens_mm": 35.0, "sensor_width_mm": 36.0, "sensor_height_mm": 24.0,
  "sensor_fit": "AUTO", "sensor_fit_resolved": "HORIZONTAL",
  "shift_x": 0.0, "shift_y": 0.0,
  "pixel_aspect_x": 1.0, "pixel_aspect_y": 1.0,
  "clip_start": 0.1, "clip_end": 1000.0
}

This matters: the scenes do not all use a default camera. 23 scenes carry a non-default sensor size, a lens shift, or a non-perspective projection.

Camera typeViews
PERSP6,517
ORTHO323
PANO114

ORTHO views need an orthographic projection using ortho_scale, not K. PANO views are not a pinhole camera at all — 3 scenes (0062_..._stairs, 0106_..._modern-bedroom-suite, 0164_..._dog-) use a panoramic camera and cannot be reprojected with K. Filter on intrinsics.type before using the camera model.

Extrinsics are camera_matrix_world (camera-to-world, Blender convention: the camera looks down its own −Z, with +Y up).

scene_result.json → render_camera records the camera the frames were actually rendered with (type, sensor size and fit, shift, ortho_scale, clip_start, read from the scene's render camera) plus any keyframed camera property at the render frame. Where it disagrees with the intrinsics above, render_camera is right: in 2 scenes the shipped intrinsics do not describe the rendered pixels (0200_..._render-scene: lens_mm_effective 24 mm against a stored 35 mm; 0077_..._stone-gazebo-greek-style: 36 mm sensor against a recorded 100 mm).

Projecting a world point with a PERSP camera:

python
import numpy as np

M = np.array(cam["camera_matrix_world"])       # camera-to-world
K = np.array(cam["intrinsics"]["K"])
p_cam = (p_world - M[:3, 3]) @ M[:3, :3]       # world -> camera
z = -p_cam[2]                                  # Blender looks down -Z
u = K[0, 0] * (p_cam[0] / z) + K[0, 2]
v = -K[1, 1] * (p_cam[1] / z) + K[1, 2]        # image v grows downward

Object captions

metadata/object_captions.json maps asset_base_id → caption record, generated with Qwen3-VL-8B-Instruct from each asset's turntable render and its BlenderKit thumbnail (both, because a render can occlude the object and a thumbnail can flatter it), plus the BlenderKit category and tags.

json
{
  "asset_base_id": "c674300a-...",
  "name": "Floor Lamp",
  "short_caption": "black metal floor lamp",
  "caption": "A tall, slender floor lamp with a black metal pole and circular base, topped with a cylindrical off-white fabric shade.",
  "placement_prompt": "a black metal floor lamp, about 0.4 m wide and 1.6 m tall",
  "object_type": "floor lamp",
  "color_primary": "black", "material_primary": "metal", "style": "modern",
  "dims_m": [...], "size_bin": "medium",
  "blenderkit_category": "...", "blenderkit_tags": [...],
  "recognizable": true, "issues": null
}

Join on asset_base_id, which also appears in moved_object_asset_base_id, added_object_asset_base_id, and in each object entry under variants.*.objects[].

Scene recovery

scene_result.json carries what is needed to rebuild a layout in Blender: the scene asset id, the seed, and per object the asset_base_id, world position, yaw, oriented bounding box (bbox_obb_world), axis-aligned box (bbox_aabb_world), dims, and size_bin — alongside the full camera set. A regenerated group records its own seed and seed rule in groups[].regen.

Distributions

Objects per group: {2: 178, 3: 290, 4: 20, 5: 13} Size bin of the added object: {small: 342, medium: 127, tiny: 28, large: 4} Rotation angle of the rotate variant (|θ|): {180°: 147, 30–60°: 76, 60–90°: 80, 90–120°: 98, 120–150°: 89}; sign of non-180° turns {+: 176, −: 167}

Both are skewed. Object count concentrates on 3 and the added object is small more often than not, because the placement budget is driven by the size of the scene's usable cluster and small objects fit more often. Sample accordingly if you need a balanced split.

Assets are reused: 211 distinct assets across 501 groups, and within one scene all groups draw from the same object set — so object identity has roughly 165 independent samples, not 501.

Intersection flags

Until v3.0 the layout pipeline's object-versus-scene collision test never ran (see Changelog), so a placed object could pass through scene furniture or a wall. Every group × variant was audited (ia_audit render-mode depth probes with instance expansion, refined against the rendered frames), and the result ships as metadata/intersect_flags.json / .csv:

  • —intersect_flag — worst visible rigid interpenetration ≥ 1 cm. Drop the group-variant when true. ge_10cm is the high-confidence subset.
  • —soft_visible_depth_m (vegetation through an object) and displacement_risk (support on a render-displaced surface) are not in intersect_flag; drop them too for a stricter filter. regenerated marks rows of groups rebuilt in v3.0 (their rows come from the final audit of the new layout).
  • —metadata/clean_groups.txt lists the 449 groups with no flag in any variant (download_update.py --clean-only group downloads exactly those).
variantflagged in v2.3flagged since v3.0≥ 10 cm since v3.0
base20444 / 50138
move21546 / 50135
add23848 / 50139
rotate20141 / 49035

Groups clean in all variants: 244 in v2.3, 449 since v3.0. Estimated precision of the flag: ~70–80% at ≥ 1 cm, ~80–88% for ge_10cm (how_to_use and estimated_precision in the JSON).

Known issues

Measured over the whole dataset, not estimated:

  • —Intersections that could not be removed: 51 of the 257 flagged groups could not be regenerated without an intersection (no clean placement within the seeds tried, or the scene state could not be verified) and keep their v2.3 data: 0013 g2, 0026 g2, 0044 g2,3, 0055 g1,2, 0062 g2, 0101 g0,1,2, 0104 g0,1,2, 0111 g1, 0112 g0,1,2,3, 0151 g0,1,2, 0153 g0,2, 0164 g0,2, 0167 g0,1,2,3, 0197 g0,1,2,3, 0204 g0,1,2,3, 0205 g3, 0213 g2, 0232 g3, 0239 g3, 0243 g1,2,3, 0246 g0,1, 0251 g0,2,3, 0257 g0,3. They stay flagged in intersect_flags.json. Regenerated groups flagged by their own final audit: 1.
  • —`rotate` is absent from 11 groups: 3 no_candidate (every candidate turn was rejected by the footprint, support or collision checks); 3 scene_nondeterministic; 3 no_base_views (no base views at all); 2 change_not_visible (the rotation stayed hidden in more than half the views). The scene_nondeterministic groups are the 3 groups of 0139_..._starbuck-coffee: that scene has a broken driver and renders in one of two states depending on the Blender process, so rotate frames could not be guaranteed to match its stored base frames. Its base orientation images are still provided.
  • —14 scenes store RGB as 16-bit RGBA PNG (2,510 frames, inherited from the scenes' own output settings; never mixed within a scene; 16-bit BBOX / orientation images: 0 / 0): 0004_..._backrooms-liminal-space-environment, 0043_..._modern-house, 0068_..._sao-pedro-restaurant, 0069_..._brick-house-glass-wall, 0080_..._kyra-s-living-room, 0082_..._green-house, 0099_..._animated-underwater-shark, 0120_..._sunset-spa-bathroom-interior, 0178_..._backdrop-tet-lunar-new-year-scene, 0192_..._modern-luxury-bathroom, 0196_..._gas-stove-flames-animation, 0217_..._realistic-abstract-flower, 0226_..._looping-balls-rolling-animation, 0247_..._simple-sci-fi-scene-setup. PIL silently reads them as 8-bit. Read them with cv2.imread(path, cv2.IMREAD_UNCHANGED) (a uint16 BGRA array; / 65535 for floats), or check the PNG header (byte 24 = bit depth) before choosing a reader.
  • —`0077_..._stone-gazebo-greek-style`: the shipped intrinsics are wrong. They record a 100 mm sensor; the frames were rendered with a 36 mm sensor (the blend holds 3 cameras). Projections with the shipped K land too close to the image centre by a factor of 100/36.
  • —`0200_..._render-scene`: the camera lens is keyframed to 24 mm in the scene, and the keyframe overrides the lens the pipeline set, so every frame was rendered at 24 mm while each view stores lens_mm 35 — this is why the scene did not reproject (≈149 px). Use render_camera.lens_mm_effective. Its box- orientation prompts use 24 mm and are aligned; its BBOX images and the per-view intrinsics are not corrected.
  • —5 scenes carry a white/grey veil (618 images in v2.3) — the darkest pixel in the frame never goes below 60/255, which is not physically possible for a correct render of those scenes. 0213_..._huge-landscape is the worst (darkest pixel 168). Detect with: per-frame minimum luminance ≥ 60 across ≥ 80% of a scene's frames.
  • —7 scenes are severely underexposed (≈1,065 images in v2.3), several recoverable with ~3 stops of gain; 2 of those (0223_..._procedural-lightning, 0216_..._etherium-wallpaper) are black-background abstract assets by design.
  • —`0259_..._stylized-medieval-flower-garden-scene` renders the environment as a uniform flat colour — objects float in a void.
  • —`0239_..._mountain-with-alpha-trees` group 03: boxes are only 10–25 px across; the annotation is unusable at that scale.
  • —1 asset (`fbbc7924-...` Blue Striped beach chair) renders at 9.2× its manifest dimensions.
  • —At least one scene (0155_..._sunset-cloudy-ocean) places objects on a water surface treated as a support plane.

What was checked and found clean: all images are 1280×720; no frame is blurred or shows depth-of-field residue; no view is a duplicate of its own base (0 move/add views with fewer than 200 changed pixels against base); all referenced files exist and decode.

Changelog

rotsinglev1 (2026-10-04)

Added the separate folder rot_single_v1/ (tag rot-single-v1): 6,000 single-object rotation pairs on a white background. The comparison-group data are unchanged from v3.1.

v3.1 (2026-10-02)

  • —All 13,710 box-orientation images re-rendered on a pure white background. v3.0 drew the boxes over the base RGB of the view faded to 35%; that faded image is gone, so a prompt shows only the boxes. Faces (red / blue / green at 50%), black edges, cameras and box geometry are unchanged — only the background pixels differ. scene_result.json → orientation_image now reads boxctrl_v2_white (background: pure white, the source image is not drawn).
  • —Repository history squashed (2026-10-02): the Hub history is now this single commit and the tags v2.3 and v3.0 were deleted. The old orientation images (v2.3 opaque boxes, v3.0 faded background) and every other file of the earlier releases that is not part of v3.1 were removed from the Hub (about 95,300 stored files, 36.7 GB), as were metadata/updates/v3.0.json / .md. Only v3.1 can be downloaded.
  • —scripts/download_update.py updates a v3.0 copy in place and reads which release a copy is from its record or README (it knows the commits and READMEs v2.3 and v3.0 were published with); the repo is public, so no login is needed.
  • —scene_result.json (orientation legend, dataset_version), metadata/dataset_summary.json and rotate_qc.txt regenerated; updates/v3.1.json (+ .md) added.

Files (against v3.0). 1 added, 13,879 modified, 2 deleted, 249,786 unchanged (about 608 MB to download): the 13,710 orientation images, 165 scene_result.json, README, scripts and metadata; deleted: the two v3.0 update notes. Every path with its sha256 is in metadata/updates/v3.1.json; scripts/download_update.py update applies it to a v3.0 copy.

Compatibility. Paths, file names and formats are unchanged; no file was added to or removed from the scene folders. The earlier releases are no longer on the Hub (v3.0 was commit 823fc9f36d).

v3.0 (2026-09-26)

Why. The layout pipeline's only object-versus-scene 3D collision test had been dead code since its first version: it built its BVH with mathutils.bvhtree.BVHTree.FromMesh, which does not exist in Blender 5.1, and the resulting AttributeError was swallowed — every placement was accepted as collision-free. The only scene constraint that worked was the support-surface test, which cannot see walls, tabletops above the anchor, thin props or instanced geometry, so objects could pass through scene furniture and walls (object-versus-object tests were working and found no interpenetration ≥ 1 cm).

What changed.

  • —206 groups in 92 scenes regenerated — every group with a flagged variant in the v2.3 audit (257 groups in 112 scenes) was laid out again with a working, fail-closed scene collider (the geometry exactly as rendered, collection / geometry-node / particle instances realised, render-time displacement counted, a canary self-test before every run; vegetation through an object is allowed and recorded as a soft contact) and a new seed, then re-rendered completely: all four variants, masks, 3D-box masks and orientation prompts, after an audit gate on the staged layout. Each group keeps its group_index and its place in the scene; the new layout can have a different number of objects and views, so files of views that no longer exist were deleted (13,822 files) and 103,314 files were added or replaced. Groups that were not flagged keep their renders and masks byte for byte (only their orientation images changed). Provenance per group in scene_result.json → groups[].regen (seed, collider, rejections, final audit gate) and per scene in metadata/regen_summary.json. 51 flagged groups could not be regenerated and keep their v2.3 data (see Known issues).
  • —All 13,710 box-orientation images re-rendered in the BoxCtrl style (semi-transparent red/blue/green faces over the faded base RGB; the background became pure white in v3.1), replacing the opaque Workbench boxes of v2.3.
  • —`render_camera` added to every scene_result.json: the camera the frames were actually rendered with (see Camera parameters).
  • —Intersection flags published: metadata/intersect_flags.json / .csv, clean_groups.txt (see Intersection flags).
  • —metadata/index.json, dataset_summary.json, rotate_qc.txt regenerated; regen_summary.json, updates/v3.0.json (+ .md) and scripts/download_update.py added.

Files. 4,718 added, 107,180 modified, 13,822 deleted, 151,768 unchanged (34.5 GB to download): 8,315 orientation images outside regenerated groups, 165 scene_result.json, 92 masks.json. (The v3.0 manifest metadata/updates/v3.0.json was removed with the history squash in v3.1.)

Compatibility. Paths and file formats are unchanged. Code that assumed 16 views per group or a fixed object count per group must read them from scene_result.json (it always had to for some groups). v2.3 was commit c00cebaa24 (no longer on the Hub).

v2.3

rotate variant (one object turned about the vertical axis, ~30% by exactly 180°) and box-orientation images (opaque Workbench boxes), with masks and QC.

Provenance

Scenes and objects are free assets from BlenderKit. Rendered with Blender 5.1 (EEVEE Next), 1280×720, 16 or 32 samples per scene (scene_config.json → rgb_samples, absent = 32 — the rotate variant and the regenerated groups reuse each scene's own setting), headless on a single GPU. Each asset keeps its asset_base_id, so licence and attribution can be resolved against BlenderKit. Redistribution of the underlying assets is governed by their individual BlenderKit licences.

3d-layout/blender-3d-layout-v2-720p · Team Ai