eternity304/procgen-renderformer
Procgen RenderFormer Dataset Procedurally generated indoor scenes with ground-truth path-traced renders and precomputed 3-slat VAE latents, built for training RenderFormer-style neural renderers. Each sample is one scene observed from 14 camera poses along an orbit. Configs Config Scenes Samples (scene x frame) Notes main ~307,000 ~4.3 M primary training set zoom ~84,000 ~1.2 M tighter framing variant validation ~1,000 ~14 K held-out assets, not… See the full description on the dataset page: https://huggingface.co/datasets/eternity304/procgen-renderformer.
Procgen RenderFormer Dataset
Procedurally generated indoor scenes with ground-truth path-traced renders and precomputed 3-slat VAE latents, built for training [RenderFormer]-style neural renderers. Each sample is one scene observed from 14 camera poses along an orbit.
Configs
Sample layout
This is a [WebDataset]. Each sample carries:
<key>.000.webp .. <key>.013.webp GT renders, 512x512 uint8 sRGB, lossless WebP
<key>.latents.safetensors coords + per-role feats for all 14 frames
<key>.json scene provenance, camera, lighting, materialsThe latents hold three roles sharing one coordinate grid per frame (a unified-proxy encode):
Frames are concatenated along the token axis with a frame_offsets index rather than stored separately, so frame i is coords[frame_offsets[i]:frame_offsets[i+1]]. Token count varies slightly per frame (~2750) because light proxies are voxelized in per frame.
Loading
from datasets import load_dataset
ds = load_dataset("eternity304/procgen-renderformer", "validation", split="validation", streaming=True)
sample = next(iter(ds))Fields arrive as raw bytes. Decode them with:
import io, json
import numpy as np
from PIL import Image
from safetensors.torch import load as st_load
ROLES = ("shape", "slat1", "slat2", "slat3")
def decode(sample):
frames = np.stack([
np.array(Image.open(io.BytesIO(sample[k])).convert("RGB"))
for k in sorted(k for k in sample if k.endswith(".webp"))
]) # uint8 [14, 512, 512, 3]
tensors = st_load(sample["latents.safetensors"])
return frames, tensors, json.loads(sample["json"])
def frame_latents(tensors, i):
"""coords [N,3] int32 and {role: feats [N,32] fp16} for frame i."""
a, b = tensors["frame_offsets"][i], tensors["frame_offsets"][i + 1]
return tensors["coords"][a:b], {r: tensors[f"{r}_feats"][a:b] for r in ROLES}coords are integer latent-voxel indices in [0, latent_grid) where latent_grid = resolution // VAE_SPATIAL_DOWNSAMPLE. Map them to [-0.5, 0.5] centres at model time.
Composed 3D scenes (test)
For benchmarking other renderers, test_assets/ holds every mesh and texture the test scenes use: 1,947 Objaverse GLBs (byte-identical to allenai/objaverse) and 624 texture folders, 9.6 GB, with manifest.json listing each file's sha256 and the scenes using it. Each sample's JSON references them relatively: assets[].glb_rel under test_assets/objaverse/, scene.surfaces[*].dir_rel under test_assets/texture_zoo/. Compose them with neural-rendering's export_composed_scenes.py:
hf download eternity304/procgen-renderformer --repo-type dataset \
--include "test/*" "test_assets/*" --local-dir procgen
python scripts/objaverse/procgen_vae/export_composed_scenes.py \
--shards 'procgen/test/shard-*.tar' --assets-root procgen/test_assets --out-dir composedwhich writes, per sample:
composed/<key>/scene.glb: the static world-space scene (room + posed assets, PBR materials, embedded textures), exactly the mesh the VAE encoder voxelized. It is glTF Y-up: the scene JSON's Z-up world mapped by(x, y, z) -> (x, z, -y), which Blender's glTF importer undoes, so the per-frame cameras and point lights in the JSONtimeline(Z-up) line up with it. Only cameras and lights move between frames.composed/<key>/mitsuba/gtc.json+ meshes + textures (relative paths): the Mitsuba scene the GT frames were rendered from (render_diamond_orbit.py --scene .../gtc.json).
Provenance
Scenes compose meshes from Objaverse-LVIS with PBR materials from PolyHaven, ambientCG, and cgbookcase (all CC0). Every scene records the source asset UIDs under assets[].uid in its JSON, so individual renders remain traceable to their source meshes.
Renders are derivative works of the source meshes. Objaverse licenses are per-asset and mixed; consult assets[].uid against the Objaverse annotations for the terms covering any particular scene.
Generation
Rendered at 64 spp, 512x512, 14 frames per scene at 12 fps, Blender/OpenGL camera convention. Full generator configuration is embedded per scene under generator_config and _dataset.
[RenderFormer]: https://arxiv.org/abs/2505.21925 [WebDataset]: https://github.com/webdataset/webdataset
