build-small-hackathon/toy-room-v2
Tiny Toybox
A Gradio-hosted Three.js virtual pet room for the Build Small Hackathon.
The first vertical slice is Squeaky's Time Room: a plush elephant pet that watches a tiny physics room, talks in short lines, changes expression, and uses time powers on real objects.
Toy Room v2 is the hackathon build: a larger shared physics room with Squeaky, Fire Boy, Shark Girl, and Electraica active at the same time. The agents have draggable/drop-able balance bodies, toggleable generated GLB rig meshes, opt-in microphone hearing, generated WebAudio sound recipes, live room/agent vision panes, generated spell operations, waste/recycling interactions, persistent JSONL memories, and MiniCPM/OpenBMB-compatible model hooks.
Hosted Space:
https://build-small-hackathon-toy-room-v2.hf.space/toy-v2
Toy Room v2
V2 adds:
- four simultaneous AI toy agents in one larger room
- active-agent power dock for one-click ability tests across Squeaky, Fire Boy, Shark Girl, and Electraica
- draggable objects and draggable agents with standing/balance physics
- active-agent force dock for lift, toss, spin, drop, and upright-settle ragdoll-style input
- force-aware rescue behavior: toss, spin, lift, or drop an agent and nearby agents visibly move in, speak, and comfort them
- toggleable generated GLB rig meshes loaded into the live room for all four agents
- waste objects, a recycle bin, food, books, chairs, lamps, plants, balls, blocks, dominos, and a ramp
- scored recycling challenge: drag recyclable waste into the bin or let Electraica sort it during the judge demo
- generic spell ops: impulse, freeze, scale, attract, particles, lights, and pet nudges
- vision-grounded decisions: "what do you see" prompts choose an action from the active agent's camera/detected-object payload
- council vision scan: the judge demo asks all four agents to inspect the room from their own agent-view cameras
- agent vision board: all four agents continuously expose their closest perceived objects and next local action affordance
- low-level motor loop: agents execute small local perception-driven moves between slower policy calls, making the AI loop feel embodied on video
- generated object recipes: prompts such as "wish for a tiny piano" can create new physical toys from simple parts
- browser speech-synthesis talkback with per-agent voice profiles plus procedural WebAudio effects
- generated sound recipes: the model can emit bounded oscillator tones for a new spell, object, or heard sound
- opt-in microphone hearing: agents receive structured sound-input summaries and can react to loud room audio
- visible learning loop: players can teach durable rules or terms, then the runtime stack shows when a remembered lesson is used
- trace-to-training export:
/api/training-datasetsummarizes valid action traces and/api/training-dataset?format=jsonl&limit=200emits a compact MiniCPM/PET action-policy SFT JSONL pack - trace-retrieval fallback policy: when no live MiniCPM endpoint is configured, the backend first retrieves a similar validated action trace before falling back to hand-written heuristics
- partner play and reciprocal dialogue: fallback and model actions can name another agent for talk, play, share, comfort, or gather interactions, and the partner visibly answers back
- physical charades: agents receive detected stacks, lines, huddles, and wished toys from the physics scene and can guess what the player built
- one-button judge demo that teaches a rule, uses that remembered lesson, makes all four agents inspect what they see, drops an agent to trigger rescue behavior, generates an object, triggers partner play, solves a physical charade, and recycles waste through the live action loop
- live judge scorecard plus
/api/judge-statusreadiness endpoint that reports hosting, assets, MiniCPM/trace-policy status, SFT traces, runtime demo proof, and remaining optional endpoint warnings - in-room Brain Trace plus runtime stack chips for text, vision, sound, learning, trace training readiness, council scans, reciprocal dialogue, force, memory, action JSON, model status, and rig readiness
- persistent runtime memories at
data/memories/toy-room-v2.jsonl - action traces at
data/traces/pet-actions.jsonl - optional MiniCPM5 text-policy and MiniCPM-V 4.6 vision endpoints
- a Docker-backed Hugging Face Space that still serves a Gradio-mounted FastAPI app
Run Locally
This project uses uv for Python dependency management. The local virtual environment is pinned to Python 3.12 via .python-version.
Current local setup was verified with:
uv 0.11.2Python 3.12.13
./start.shstart.sh stops previous Tiny Toybox app.py processes from this workspace, starts a fresh server, and prints the active URLs. Open http://localhost:65372 for the page directory.
Useful local URLs:
- Page directory:
http://localhost:65372/pages - Toy Room v2:
http://localhost:65372/toy-v2 - Toy room:
http://localhost:65372/toy - Procedural model lab:
http://localhost:65372/models - Blender rig previews and GLBs:
http://localhost:65372/blender-models - Layered part concept refs:
http://localhost:65372/parts-lab - Fire Boy rigged viewer:
http://localhost:65372/fireboy-rigged
The default app port is 65372 to avoid common local preview conflicts. To choose another port for a one-off run:
PORT=65400 ./start.shTo stop it:
./shutdown.shTo restart everything from this project, run:
./start.shManual uv flow:
uv sync --python 3.12
.venv/bin/python app.pystart.sh uses uv sync first, then runs the uv-created .venv/bin/python directly so shutdown.sh can stop the app cleanly by PID.
Blender And SAM Character Assets
Blender is expected on PATH as blender. On this machine that is a wrapper in ~/.local/bin/blender pointing at /Applications/Blender.app/Contents/MacOS/Blender.
Regenerate all character assets, rig previews, beauty renders, object lineups, GLBs, and the contact sheet with:
./scripts/render_blender_models.shOutputs are written to:
assets/generated/rigged/*.glbassets/generated/previews/*.png
Clean raw fal/SAM GLB extractions from potential-char-images/extracted-from-sam with:
./scripts/clean_sam_models.shCleaned SAM outputs are written to:
assets/generated/sam-cleaned/*.glbassets/generated/sam-standing-rigged/*.glbassets/generated/previews/*-sam-cleaned.pngassets/generated/previews/*-sam-standing-*.png
Layered 2D part concept outputs are written to:
assets/generated/part-concepts/*-parts-sheet.pngassets/generated/part-concepts/individual/*/*.pngfor the original sheet-derived v1 cropsassets/generated/part-concepts/individual-v2/*/*.pngfor the cleaner individually generated v2 refsassets/generated/part-concepts/*-individual-v2-contact.pngassets/generated/part-concepts/parts-individual-v2-contact.pngassets/generated/part-concepts/parts-manifest.json
The v2 refs are the better input set for fal/SAM object extraction because each base body or prop is generated as one isolated image. The four base bodies are standing, while props stay separate for later Blender bone/socket attachment.
Generate fal SAM 3D Object GLBs from the four local source images with:
/Library/Frameworks/Python.framework/Versions/3.14/bin/python3 scripts/generate_sam_3d_models.pyThat script sends images as data URLs, which avoids needing fal files upload permissions.
Generate fal SAM 3D Object GLBs from the v2 isolated base bodies, clothing, backpacks, and props with:
/Library/Frameworks/Python.framework/Versions/3.14/bin/python3 scripts/generate_sam_part_models.pyPart-level SAM outputs are written to:
assets/generated/part-models/raw/*/*-sam.glbassets/generated/part-models/raw/*/*-sam-result.jsonassets/generated/part-models/sam-part-inputs.json
You can also run a focused pass, for example:
/Library/Frameworks/Python.framework/Versions/3.14/bin/python3 scripts/generate_sam_part_models.py --bases-only
/Library/Frameworks/Python.framework/Versions/3.14/bin/python3 scripts/generate_sam_part_models.py fire-boy-fluteRig the four v2 standing base bodies and build socketed assembly test GLBs with:
./scripts/rig_part_base_models.shThe rig/assembly pass writes:
assets/generated/part-models/rigged-bases/*-base-rigged.glbassets/generated/part-models/assemblies/*-assembled.glbassets/generated/part-models/mixamo-fbx/*-base-mesh.fbxfor Mixamo auto-rig upload testsassets/generated/part-models/mixamo-fbx/*-base-rigged.fbxfor rigged FBX inspectionassets/generated/part-models/blend-scenes/*-assembly.blendassets/generated/previews/*-part-base-rigged.pngassets/generated/previews/*-part-assembly.png
Optional Local Model Hook
The app can use a local OpenAI-compatible PET LLM endpoint. MiniCPM5 local mode is the recommended first text-policy brain:
scripts/start_with_minicpm5.shThat script uses Ollama and hf.co/openbmb/MiniCPM5-1B-GGUF:Q4_K_M.
Manual PET LLM flow:
scripts/pull_minicpm5_ollama.sh
export TOYBOX_LLM_ENDPOINT=http://127.0.0.1:11434/v1/chat/completions
export TOYBOX_LLM_MODEL=hf.co/openbmb/MiniCPM5-1B-GGUF:Q4_K_M
./start.shCheck the model endpoint:
uv run python scripts/check_pet_llm.pyOptional Hosted Model Hook
The hosted Space can call any OpenAI-compatible chat-completions endpoint. For Hugging Face Inference Providers, set these Space variables/secrets:
hf spaces variables add build-small-hackathon/toy-room-v2 \
-e TOYBOX_LLM_ENDPOINT=https://router.huggingface.co/v1/chat/completions \
-e TOYBOX_LLM_MODEL=provider-backed/chat-model-id
hf spaces secrets add build-small-hackathon/toy-room-v2 \
-s TOYBOX_LLM_API_KEYOptional org billing header:
hf spaces variables add build-small-hackathon/toy-room-v2 \
-e TOYBOX_LLM_BILL_TO=your-hf-org-or-usernameTOYBOX_LLM_API_KEY may also be supplied as HF_TOKEN for Hugging Face endpoints, or OPENAI_API_KEY for OpenAI endpoints. The /api/model-status endpoint reports whether a hosted endpoint is active, configured but missing a secret, or falling back.
RunPod serverless endpoints are also supported when they expose an OpenAI-compatible chat-completions route:
hf spaces variables add build-small-hackathon/toy-room-v2 \
-e TOYBOX_LLM_ENDPOINT=https://api.runpod.ai/v2/YOUR_ENDPOINT_ID/openai/v1/chat/completions \
-e TOYBOX_LLM_MODEL=openbmb/MiniCPM5-1B-or-your-served-model-id
hf spaces secrets add build-small-hackathon/toy-room-v2 \
-s RUNPOD_API_KEYFor a RunPod MiniCPM-V visual cortex, set TOYBOX_VISION_ENDPOINT and TOYBOX_VISION_MODEL to the corresponding OpenAI-compatible vision endpoint/model. The same RUNPOD_API_KEY secret is reused unless TOYBOX_VISION_API_KEY is supplied. /api/model-status reports provider: runpod and mode: runpod-openai-compatible for these endpoints.
If no endpoint is configured, the public build uses a deterministic heuristic fallback so the game stays playable. If an endpoint is configured but unavailable, the pet enters visible asleep/model-off mode by default. Set TOYBOX_ALLOW_HEURISTIC_FALLBACK=1 only for local debugging when you want heuristic behavior even after a model endpoint fails.
The current OpenBMB/MiniCPM path is local-first through Ollama because the public HF router metadata did not expose provider-backed OpenBMB MiniCPM chat models during this build. The game still uses the same action JSON contract, so a hosted MiniCPM endpoint can be connected by setting TOYBOX_LLM_ENDPOINT, TOYBOX_LLM_MODEL, and a secret token.
Check Modal remote execution:
uv run --with modal modal run scripts/modal_square_smoke.pyMeasure the current local runtime:
uv run python scripts/measure_runtime.py --samples 5On macOS, power sampling needs sudo. If you already have a cached sudo session:
uv run python scripts/measure_runtime.py --samples 5 --powerMiniCPM-V 4.6 can be added as the pet's visual cortex. It reads the rendered room camera frame and returns perception plus face blendshape hints, while MiniCPM5 remains the faster action/personality model:
scripts/start_with_minicpmv46_vision.shThat script uses Ollama models:
hf.co/openbmb/MiniCPM5-1B-GGUF:Q4_K_Mfor PET-LLM actionsopenbmb/minicpm-v4.6for vision perception
MiniCPM-V 4.6 local vision currently needs Ollama 0.30.0 or newer. The script checks this before pulling the vision model.
Check only the vision endpoint:
TOYBOX_VISION_ENDPOINT=http://127.0.0.1:11434/api/chat \
TOYBOX_VISION_MODEL=openbmb/minicpm-v4.6 \
uv run python scripts/check_vision_endpoint.pyIf no endpoint is configured, the app uses a deterministic fallback policy so the toy remains playable. If an endpoint is configured but cannot be used, the default behavior is visible asleep/model-off mode rather than silently pretending a heuristic is the model.
Action traces are written to data/traces/pet-actions.jsonl by default. These become the seed dataset for a later distilled pet-policy model.
Code Shape
src/pet_policy.pyis the small orchestration layer.src/model_policy.pytalks to text/PET-LLM endpoints.src/vision_policy.pytalks to MiniCPM-V-style image endpoints.src/pet_actions.pyvalidates actions, face blendshapes, powers, and fallback behavior.objectRecipein pet actions is the bounded generated-content path for wishable physical toys.src/pet_payload.pyowns scene compaction, target selection, and touch detection.src/pet_payload.pyalso detects physical arrangements so model/fallback policies can ground guesses in object positions.frontend/toybox/pet.jsowns character meshes, face drawing, and blendshape interpolation.frontend/toybox/pet_balance.jsowns the hidden weighted standing/balance physics rig.frontend/toybox/senses.jsowns user-view, pet-view, audio, and balance feeds.frontend/toybox/room.jsowns the room shell, physics objects, and history.frontend/toybox/powers.jsowns executable pet powers and target selection.
See docs/modal-1bit-model-plan.md for the current Modal, MiniCPM-V, MiniCPM5, and 1-bit policy plan.
