Skip to content

← ArchivePaper2022

MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, Ziwei Liu

arXivAcademic1051 citesMotion Synthesis

First diffusion-based text-driven motion framework enabling fine-grained body-part control and arbitrary-length synthesis.

Abstract

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions conditioned on natural languages. However, it remains challenging to achieve diverse and fine-grained motion generation with various text inputs. To address this problem, we propose MotionDiffuse, one of the first diffusion model-based text-driven motion generation frameworks, which demonstrates several desired properties over existing methods. 1) Probabilistic Mapping. Instead of a deterministic language-motion mapping, MotionDiffuse generates motions through a series of denoising steps in which variations are injected. 2) Realistic Synthesis. MotionDiffuse excels at modeling complicated data distribution and generating vivid motion sequences. 3) Multi-Level Manipulation. MotionDiffuse responds to fine-grained instructions on body parts, and arbitrary-length motion synthesis with time-varied text prompts. Our experiments show MotionDiffuse outperforms existing SoTA methods by convincing margins on text-driven motion generation and action-conditioned motion generation. A qualitative analysis further demonstrates MotionDiffuse's controllability for comprehensive motion generation.

How to read this

Category
Method: diffusion-based text-driven motion generation
Contributions
  • One of the first diffusion-model frameworks for text-driven human motion, giving a probabilistic (non-deterministic) language-to-motion mapping
  • Models complex motion distributions for realistic synthesis
  • Supports multi-level manipulation: fine-grained body-part instructions and arbitrary-length synthesis with time-varied text prompts
Context
Builds on text-to-motion learned from paired datasets such as Guo et al. HumanML3D (2022), bringing the denoising-diffusion paradigm to the motion domain.Builds on: Generating Diverse and Natural 3D Human Motions from Text
Correctness
Reports outperforming prior state of the art on text-driven and action-conditioned generation by the paper's account; as a generative diffusion model it produces diverse plausible motions rather than unique ground truth, and quality depends on the text-motion training data.
Clarity
Accessible if you know diffusion models; a first pass conveys the properties, a second pass covers the denoising formulation and the body-part/time-varied conditioning.
How to read it
First pass for the three claimed properties (probabilistic, realistic, multi-level); second pass on the conditioned denoising and how body-part and time-varied text control are injected.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →