← ArchivePaper2022
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, Ziwei Liu
First diffusion-based text-driven motion framework enabling fine-grained body-part control and arbitrary-length synthesis.
Abstract
Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions conditioned on natural languages. However, it remains challenging to achieve diverse and fine-grained motion generation with various text inputs. To address this problem, we propose MotionDiffuse, one of the first diffusion model-based text-driven motion generation frameworks, which demonstrates several desired properties over existing methods. 1) Probabilistic Mapping. Instead of a deterministic language-motion mapping, MotionDiffuse generates motions through a series of denoising steps in which variations are injected. 2) Realistic Synthesis. MotionDiffuse excels at modeling complicated data distribution and generating vivid motion sequences. 3) Multi-Level Manipulation. MotionDiffuse responds to fine-grained instructions on body parts, and arbitrary-length motion synthesis with time-varied text prompts. Our experiments show MotionDiffuse outperforms existing SoTA methods by convincing margins on text-driven motion generation and action-conditioned motion generation. A qualitative analysis further demonstrates MotionDiffuse's controllability for comprehensive motion generation.
How to read this
- Category
- Method: diffusion-based text-driven motion generation
- Contributions
- One of the first diffusion-model frameworks for text-driven human motion, giving a probabilistic (non-deterministic) language-to-motion mapping
- Models complex motion distributions for realistic synthesis
- Supports multi-level manipulation: fine-grained body-part instructions and arbitrary-length synthesis with time-varied text prompts
- Context
- Builds on text-to-motion learned from paired datasets such as Guo et al. HumanML3D (2022), bringing the denoising-diffusion paradigm to the motion domain.Builds on: Generating Diverse and Natural 3D Human Motions from Text
- Correctness
- Reports outperforming prior state of the art on text-driven and action-conditioned generation by the paper's account; as a generative diffusion model it produces diverse plausible motions rather than unique ground truth, and quality depends on the text-motion training data.
- Clarity
- Accessible if you know diffusion models; a first pass conveys the properties, a second pass covers the denoising formulation and the body-part/time-varied conditioning.
- How to read it
- First pass for the three claimed properties (probabilistic, realistic, multi-level); second pass on the conditioned denoising and how body-part and time-varied text control are injected.
Built upon by
Nothing yet.
Related work
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →