Skip to content

← ArchivePaper2018

DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills

Xue Bin Peng, Pieter Abbeel, Sergey Levine, Michiel van de Panne

SIGGRAPHAcademic600 cites34 descendantsMotion Synthesis

Physics-based character controller using deep RL to imitate reference motion clips, achieving diverse athletic skills with natural-looking dynamics.

Abstract

A longstanding goal in character animation is to combine data-driven specification of behavior with a system that can execute a similar behavior in a physical simulation, enabling realistic responses to perturbations and environmental variation. We show that well-known reinforcement learning methods can be adapted to learn robust control policies capable of imitating a broad range of example motion clips, while also learning complex recoveries, adapting to changes in morphology, and accomplishing user-specified goals. The method handles keyframed motions, highly dynamic actions such as motion-captured flips and spins, and retargeted motions. By combining a motion-imitation objective with a task objective, characters can be trained to react intelligently in interactive settings.

How to read this

Category
Method: physics-based character control via deep reinforcement learning
Contributions
  • Adapts deep RL to learn control policies that imitate a broad range of reference motion clips in physics simulation
  • Combines a motion-imitation objective with a task objective so characters pursue user goals while staying natural
  • Handles keyframed, highly dynamic (flips, spins), and retargeted motions, and learns complex recoveries plus adaptation to morphology changes
Context
Builds on the authors' earlier deep-RL locomotion work (Peng et al., Terrain-Adaptive Locomotion Skills Using Deep Reinforcement Learning), advancing example-guided imitation of reference motion.Builds on: Terrain-Adaptive Locomotion Skills Using Deep Reinforcement Learning
Correctness
Demonstrated across diverse athletic skills with natural-looking dynamics and perturbation recovery; results depend on quality reference clips and per-skill training, and the imitation-plus-task objective balance is central to the behavior obtained.
Clarity
Accessible in motivation; a first pass conveys the imitation-plus-task idea, a second pass is needed for the reward design and training setup.
How to read it
Focus on the reward formulation (imitation plus task) and how reference clips are used; a second pass pays off for the RL training details if you plan to reproduce or extend it.

Builds on

Built upon by

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →