← ArchivePaper2026
MotionBricks: Scalable Real-Time Motions with Modular Latent Generative Model and Smart Primitives
Tingwu Wang, Olivier Dionne, Michael De Ruyter, David Minor, Davis Rempe, Kaifeng Zhao, Mathis Petrovich, Ye Yuan, Chenran Li, Zhengyi Luo, Brian Robison, Xavier Blackwell, Bernardo Antoniazzi, Xue Bin Peng, Yuke Zhu, Simon Yuen
MotionBricks is a real-time generative motion framework: a modular latent backbone models over 350,000 clips in one network and a smart-primitive interface authors navigation and object interaction plug-and-play, reaching 15,000 FPS at 2ms latency and deployed zero-shot to a Unitree G1 humanoid.
Abstract
Despite transformative advances in generative motion synthesis, real-time interactive motion control remains dominated by traditional techniques. In this work, we identify two key challenges in bridging research and production: 1) Real-time scalability: Industry applications demand real-time generation of a vast repertoire of motion skills, while generative methods exhibit significant degradation in quality and scalability under real-time computation constraints, and 2) Integration: Industry applications demand fine-grained multi-modal control involving velocity commands, style selection, and precise keyframes, a need largely unmet by existing text- or tag-driven models. Moreover, a systematic motion design interface for generative models remains absent. To overcome these limitations, we introduce MotionBricks: a large-scale, real-time generative framework with a two-fold solution. First, we propose a large-scale modular latent generative backbone tailored for robust real-time motion generation, effectively modeling a dataset of over 350,000 motion clips with a single model. Second, we introduce smart primitives that provide a unified, robust, and intuitive interface for authoring both navigation and object interaction. Notably, MotionBricks applies to new downstream tasks in a zero-shot manner, where no finetuning or task-specific tagging is required. Applications can be designed in a plug-and-play manner like assembling bricks without expert animation knowledge, enabling an accessible interface for applications in animation and robotics. Quantitatively, we show that MotionBricks produces state-of-the-art motion quality on open-source and proprietary datasets of various scales, while also achieving a real-time throughput of 15,000 FPS with 2ms latency. We demonstrate the flexibility and robustness of MotionBricks in a complete production-level animation demo, covering navigation and object-scene interaction across various styles with a unified model. To showcase our framework's application beyond animation, we deploy MotionBricks on the Unitree G1 humanoid robot to demonstrate its flexibility and generalization for real-time robotic control.
How to read this
- Category
- Method / system: a real-time generative motion framework (modular latent backbone plus a smart-primitive authoring interface) for animation and robotics
- Contributions
- A modular latent generative backbone built on motion in-betweening that models a single dataset of over 350,000 motion clips and runs far above real time (reported 15,000 FPS at 2ms latency).
- Smart primitives (smart locomotion and smart object) that give a unified, plug-and-play interface for authoring navigation and object interaction without animation graphs or per-task tagging, and that apply zero-shot to new downstream tasks.
- An end-to-end production-level demonstration spanning a UE5 animation demo (locomotion, acrobatics, object-scene interaction) and real-world whole-body control deployed on a Unitree G1 humanoid robot.
- Context
- A generative successor to control-driven runtime motion methods, from motion matching and phase-functioned or learned controllers to physics-based RL and diffusion motion models, aimed squarely at the real-time, controllable, production-scale regime those earlier lines could not satisfy all at once.Builds on: Motion Matching and The Road to Next-Gen Animation · Phase-Functioned Neural Networks for Character Control · Learned Motion Matching · DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills · Human Motion Diffusion Model · DeepPhase: Periodic Autoencoders for Learning Motion Phase Manifolds
- Correctness
- Quality is reported as state of the art on open-source and proprietary datasets of various scales; the headline 15,000 FPS and 2ms numbers are throughput and latency for the neural backbone under the authors' own setup, and the robot and UE5 results are qualitative production demonstrations rather than controlled user studies, so generality beyond the shown scenarios should be read with that in mind.
- Clarity
- The two framed challenges (real-time scalability and fine-grained integration) and the bricks metaphor make the high-level design easy to follow; the structured latent backbone and the in-betweening formulation are the parts that reward a careful second read.
- How to read it
- First pass for the two production challenges and the smart-primitive interface idea; second pass on the structured modular latent design and the in-betweening backbone if you build, extend, or benchmark real-time motion-control systems.
Builds on
- Motion Matching and The Road to Next-Gen Animation 2016
- Phase-Functioned Neural Networks for Character Control 2017
- Learned Motion Matching 2020
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills 2018
- Human Motion Diffusion Model 2022
- DeepPhase: Periodic Autoencoders for Learning Motion Phase Manifolds 2022
Built upon by
Nothing yet.
Related work
- Character Motion Synthesis by Topology Coordinates 2009 / CGF
- Mode-Adaptive Neural Networks for Quadruped Motion Control 2018 / SIGGRAPH
- Motion Retargeting for Crowd Simulation 2015 / DigiPro
- SMPL: A Skinned Multi-Person Linear Model 2015 / SIGGRAPH Asia
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →