← ArchivePaper2026
Two2Four: Generative Quadruped Puppeteering from Human Motion
Fatemeh Zargarbashi, Zehong Qiu, Dhruv Agrawal, Stelian Coros, Robert Sumner, Martin Guay, Jakob Buhmann
Two stage generative diffusion model trained purely on quadruped motion converts ordinary human motion into plausible controllable quadruped puppeteering with limb level control.
How to read this
- Category
- generative cross morphology motion transfer paper
- Contributions
- Trains a two stage diffusion model purely on quadruped motion data, avoiding the need for paired human to quadruped motion, and uses it to convert ordinary human motion into controllable quadruped motion
- Separates control signals into precise features (travel direction, target location) applied through direct network conditioning and imprecise features (limb positions, root height) applied through inference time inpainting, matching the ambiguity inherent in cross morphology transfer
- Supports fine grained control including head direction and individual limb puppeteering, plus semantic actions such as sitting, lying and jumping obtained by inpainting features like root height
- Splits generation into a trajectory stage over a subset of joints and a full body stage, which the authors show improves alignment and makes inpainting more effective when training data is limited, with autoregressive rollout for long horizon motion
- Reports improved realism and controllability, including reduced foot sliding, compared to existing retargeting baselines such as VQ-PAE
- Context
- Two2Four builds on the generative diffusion motion lineage established by MDM and EDGE, and responds directly to prior cross morphology retargeting work such as VQ-PAE and optimisation based puppeteering, which either lack paired data or lack semantic alignment. Its central move is to sidestep the missing paired human animal data problem entirely by training only on quadruped motion and mapping human input in at inference time through conditioning and inpainting rather than a jointly learned embedding.
- Correctness
- The claims rest on comparisons against baseline retargeting and data driven methods plus qualitative demonstrations across walking, running, jumping, sitting and lying, with the specific quantitative results held in the paper's experiments section rather than in the introduction read here. The quadruped training data itself is explicitly described as small scale (in the tradition of the MANN dataset), which is the reason the two stage and inpainting design exists, but it also means generalisation to gaits or behaviours far outside that training distribution remains a real risk.
- Clarity
- Written for a graphics research audience familiar with diffusion models, but the central idea, splitting control into precise conditioning versus imprecise inpainting, is explainable to a technical animator without deep diffusion background.
- How to read it
- First pass, read the abstract and study Figure 1 and Figure 2 to see the two stage pipeline and the distinction between conditioning and inpainting control. Second pass, read the introduction's discussion of precise versus imprecise features closely, since that framing is the paper's core technical bet. Third pass, read the results section for the quantitative comparison against baselines like VQ-PAE and the foot sliding metric, which is the practical evidence for anyone considering this for a production puppeteering pipeline.
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →