binhpham/reachy-mini-motion-planner-27b
reachy-mini-motion-planner-27b
Text-to-motion planner for Reachy Mini: a prompt goes in, and a ready-to-play head, antenna and body trajectory comes out. This repo is a complete serving bundle:
python -m inference.server --bundle binhpham/reachy-mini-motion-planner-27b # from the reachy-motion-generator project
curl -s localhost:8000/generate-dense -H 'content-type: application/json' \
-d '{"prompt": "sneezing. You build up and then sneeze loudly.", "n": 2}'
# all three planners in one process (one GPU, shared generator), picked per request with "effort"
python -m inference.server --bundle high=binhpham/reachy-mini-motion-planner-27b \
--bundle medium=binhpham/reachy-mini-motion-planner-4b --bundle low=binhpham/reachy-mini-motion-planner-0.8b
curl -s localhost:8000/generate-dense -H 'content-type: application/json' -d '{"prompt": "startled. A door slams.", "effort": "medium"}'The response holds the recipe, a one-line idea, and moves: Reachy Mini recorded-move dicts ({"time", "set_target_data": [{"head": 4x4, "antennas", "body_yaw"}]}), already projected onto the robot's reachable set. Serving effort: "high" runs this model. Speed: ~0.84 s median per prompt on one RTX PRO 6000 (FP8 + fine-tuned MTP speculative decoding; planner 0.75 s). GPU memory: ~40 GB with FP8.
How the planner is prompted
Compact system prompt (units, recipe grammar, 4 motion rules, 3 examples), user message = the prompt (word. one sentence of context. works best). The answer is {"idea", "recipe"} with thinking off. Allow at least 400 output tokens.
Training
- Data: 5,872 teacher rows: hand-written recipes (Claude), build-up/release events (×3), 287 seeds, and 5,000 scenarios (Astra) re-authored in a lively style by Codex (gpt-6-astra). Each prompt is also trained as its bare word and its sentence alone. Rows near any evaluation prompt are removed (embedding filter plus a keyword blocklist for the out-of-distribution probes). The data is published as binhpham/reachy-mini-massive-motion-library.
- Training: LoRA r = 32 on the attention and MLP projections, loss on the answer only, best checkpoint by held-out loss. The MTP head was then fine-tuned on the planner's own outputs, which makes speculative decoding faster without changing outputs.
Evaluation (prompts never seen in training, 12 samples per probe)
Try it in the browser: binhpham/reachy-mini-motion-generator.
