Team Ai
Datasetpublic

binhpham/reachy-mini-motion-synth

Reachy Mini Text-to-Motion (real + synthetic) Expressive motions for Reachy Mini: each clip pairs a text prompt (an emotion, reaction, character or situation) with a head / antenna / body-yaw trajectory. Built to train and evaluate text → motion models that generalise beyond the handful of real recordings. source clips duration real: Pollen emotions library, reachability-fixed 85 8.6 min generated_batch1: 154 prompts × 4 plan variants × 2 render seeds 1,232 98 min… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/reachy-mini-motion-synth.

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
0likes145downloads
Dataset Card

Reachy Mini Text-to-Motion (real + synthetic)

Expressive motions for Reachy Mini: each clip pairs a text prompt (an emotion, reaction, character or situation) with a head / antenna / body-yaw trajectory. Built to train and evaluate text → motion models that generalise beyond the handful of real recordings.

sourceclipsduration
real: Pollen emotions library, reachability-fixed858.6 min
generated_batch1: 154 prompts × 4 plan variants × 2 render seeds1,23298 min
generated_batch2: 133 prompts × 6 plan variants × 1 render seed79868 min
total2,115~175 min

Files

  • —metadata.jsonl: one row per clip: file, source, name, word, prompt, duration_s, and for generated clips variant, mirrored, render_seed. Real clips carry heldout_eval: true for the 12 emotions held out of all training.
  • —data/real/*.json, data/generated/batch{1,2}/*.json: the motions, in the same format as the Pollen libraries (description, time [s], set_target_data[i] = {head: 4×4 pose, antennas: [right, left] rad, body_yaw: rad}). They are playable with the Reachy Mini SDK's recorded-move player and stored at 25 Hz. Every frame is IK-reachable.
  • —plans.jsonl: the motion plan each generated clip was rendered from: keyframes every 0.5 s of ear droop R/L (deg; 0 = up), head pitch/roll/yaw (deg; +pitch = head down), head height (mm), body yaw (deg), and energy (RMS of fast motion detail, deg).
  • —recipes.json: the hand-written motion recipe (a small keyframe DSL) behind each generated prompt.
  • —paraphrases.json: up to 4 paraphrased captions per prompt (LLM-written, leak-filtered), for caption augmentation.
  • —eval_prompts.txt: the prompts used for evaluation (12 held-out real emotions and 18 out-of-distribution prompts). No generated prompt or paraphrase is within cosine 0.72 of any of them (Qwen3-Embedding-0.6B, last-token pooling).

How the synthetic clips were made

  1. 1.Planner. An LLM (Claude) wrote a motion recipe for each prompt. Each recipe was expanded into several randomised plans: amplitude ×0.75–1.25, tempo ×0.8–1.25, ear asymmetry, and a sagittal mirror on odd variants. Plans are low-passed at 1 Hz.
  2. 2.Renderer. A 22M-parameter flow-matching transformer, trained only on real motion (the emotions and dances libraries, minus the held-out clips), turns plan → full 25 Hz motion and adds the fast, organic detail the plan leaves out. From true plans it reproduces held-out clips well enough to be identified among 12 held-out emotions 89% of the time.
  3. 3.Projection. Every frame is projected onto the robot's reachable set using the SDK's analytical IK.

The synthetic clips are model outputs, not recordings. They reflect one LLM's idea of how each prompt should move, filtered through a renderer trained on ~9 minutes of real motion, so treat them as a teacher distribution, not ground truth.

Credits

Real clips are derived from pollen-robotics/reachy-mini-emotions-library (Apache-2.0). Frames that were unreachable (20 of 85 clips, 1,232 of 25,824 frames) were projected to the nearest reachable pose.