Team Ai
Datasetpublic

eternity304/procgen-renderformer

Procgen RenderFormer Dataset Procedurally generated indoor scenes with ground-truth path-traced renders and precomputed 3-slat VAE latents, built for training RenderFormer-style neural renderers. Each sample is one scene observed from 14 camera poses along an orbit. Configs Config Scenes Samples (scene x frame) Notes main ~307,000 ~4.3 M primary training set zoom ~84,000 ~1.2 M tighter framing variant validation ~1,000 ~14 K held-out assets, not… See the full description on the dataset page: https://huggingface.co/datasets/eternity304/procgen-renderformer.

sourceHugging Facecc-by-4.0updated 12h agoView on Hugging Face
0likes638downloads
Dataset Card

Procgen RenderFormer Dataset

Procedurally generated indoor scenes with ground-truth path-traced renders and precomputed 3-slat VAE latents, built for training [RenderFormer]-style neural renderers. Each sample is one scene observed from 14 camera poses along an orbit.

Configs

ConfigScenesSamples (scene x frame)Notes
main~307,000~4.3 Mprimary training set
zoom~84,000~1.2 Mtighter framing variant
validation~1,000~14 Kheld-out assets, not just held-out scenes
caustic~20,000~62 Kcaustic dielectric dataset (6,144 SPP Mitsuba GT + 4th material slat: slat3)
caustic_envmap~35,600~180 Kenvironment-map caustic dataset (lit by 997 Poly Haven HDRIs with continuous orbit + Qwen-Image VAE latents)
test997~4 Knon-caustic point-light test set (1,024 SPP Mitsuba GT, 4 frames per scene; even split of Cornell box with a moving light, camera trajectory, dolly zoom and walkthrough; standard PBR assets only, slat3 holds the non-caustic baseline)

Sample layout

This is a [WebDataset]. Each sample carries:

<key>.000.webp .. <key>.013.webp   GT renders, 512x512 uint8 sRGB, lossless WebP
<key>.latents.safetensors          coords + per-role feats for all 14 frames
<key>.json                         scene provenance, camera, lighting, materials

The latents hold three roles sharing one coordinate grid per frame (a unified-proxy encode):

RoleMeaning
shapegeometry
slat1PBR: albedo / metallic / roughness / alpha
slat2emissive / normal / occlusion + light flag
slat3caustic: $\sigma_t$ (extinction) / IOR / $\xi$ (transmittance) / albedo (scattering)

Frames are concatenated along the token axis with a frame_offsets index rather than stored separately, so frame i is coords[frame_offsets[i]:frame_offsets[i+1]]. Token count varies slightly per frame (~2750) because light proxies are voxelized in per frame.

Loading

python
from datasets import load_dataset

ds = load_dataset("eternity304/procgen-renderformer", "validation", split="validation", streaming=True)
sample = next(iter(ds))

Fields arrive as raw bytes. Decode them with:

python
import io, json
import numpy as np
from PIL import Image
from safetensors.torch import load as st_load

ROLES = ("shape", "slat1", "slat2", "slat3")

def decode(sample):
    frames = np.stack([
        np.array(Image.open(io.BytesIO(sample[k])).convert("RGB"))
        for k in sorted(k for k in sample if k.endswith(".webp"))
    ])                                              # uint8 [14, 512, 512, 3]
    tensors = st_load(sample["latents.safetensors"])
    return frames, tensors, json.loads(sample["json"])

def frame_latents(tensors, i):
    """coords [N,3] int32 and {role: feats [N,32] fp16} for frame i."""
    a, b = tensors["frame_offsets"][i], tensors["frame_offsets"][i + 1]
    return tensors["coords"][a:b], {r: tensors[f"{r}_feats"][a:b] for r in ROLES}

coords are integer latent-voxel indices in [0, latent_grid) where latent_grid = resolution // VAE_SPATIAL_DOWNSAMPLE. Map them to [-0.5, 0.5] centres at model time.

Composed 3D scenes (test)

For benchmarking other renderers, test_assets/ holds every mesh and texture the test scenes use: 1,947 Objaverse GLBs (byte-identical to allenai/objaverse) and 624 texture folders, 9.6 GB, with manifest.json listing each file's sha256 and the scenes using it. Each sample's JSON references them relatively: assets[].glb_rel under test_assets/objaverse/, scene.surfaces[*].dir_rel under test_assets/texture_zoo/. Compose them with neural-rendering's export_composed_scenes.py:

bash
hf download eternity304/procgen-renderformer --repo-type dataset \
    --include "test/*" "test_assets/*" --local-dir procgen
python scripts/objaverse/procgen_vae/export_composed_scenes.py \
    --shards 'procgen/test/shard-*.tar' --assets-root procgen/test_assets --out-dir composed

which writes, per sample:

  • —composed/<key>/scene.glb: the static world-space scene (room + posed assets, PBR materials, embedded textures), exactly the mesh the VAE encoder voxelized. It is glTF Y-up: the scene JSON's Z-up world mapped by (x, y, z) -> (x, z, -y), which Blender's glTF importer undoes, so the per-frame cameras and point lights in the JSON timeline (Z-up) line up with it. Only cameras and lights move between frames.
  • —composed/<key>/mitsuba/gtc.json + meshes + textures (relative paths): the Mitsuba scene the GT frames were rendered from (render_diamond_orbit.py --scene .../gtc.json).

Provenance

Scenes compose meshes from Objaverse-LVIS with PBR materials from PolyHaven, ambientCG, and cgbookcase (all CC0). Every scene records the source asset UIDs under assets[].uid in its JSON, so individual renders remain traceable to their source meshes.

Renders are derivative works of the source meshes. Objaverse licenses are per-asset and mixed; consult assets[].uid against the Objaverse annotations for the terms covering any particular scene.

Generation

Rendered at 64 spp, 512x512, 14 frames per scene at 12 fps, Blender/OpenGL camera convention. Full generator configuration is embedded per scene under generator_config and _dataset.

[RenderFormer]: https://arxiv.org/abs/2505.21925 [WebDataset]: https://github.com/webdataset/webdataset

eternity304/procgen-renderformer · Team Ai