Skip to content

← ArchivePaper2011

Realtime Performance-Based Facial Animation

Thibaut Weise, Sofien Bouaziz, Hao Li, Mark Pauly

SIGGRAPHAcademic2 descendantsFacial

Real-time performance-based facial animation driven by depth camera input, enabling live facial retargeting at interactive rates.

Abstract

This paper presents a system for performance-based character animation that enables any user to control the facial expressions of a digital avatar in realtime. The user is recorded in a natural environment using a non-intrusive, commercially available 3D sensor. The simplicity of this acquisition device comes at the cost of high noise levels in the acquired data. To effectively map low-quality 2D images and 3D depth maps to realistic facial expressions, we introduce a novel face tracking algorithm that combines geometry and texture registration with pre-recorded animation priors in a single optimization. Formulated as a maximum a posteriori estimation in a reduced parameter space, our method implicitly exploits temporal coherence to stabilize the tracking. We demonstrate that compelling 3D facial dynamics can be reconstructed in realtime without the use of face markers, intrusive lighting, or complex scanning hardware. This makes our system easy to deploy and facilitates a range of new applications, e.g. in digital gameplay or social interactions.

How to read this

Category
System / method: real-time performance-based facial animation from a depth sensor
Contributions
  • A system letting any user drive a digital avatar's facial expressions in real time, recorded with a non-intrusive commercial 3D sensor in a natural environment
  • A face tracking algorithm combining geometry and texture registration with pre-recorded animation priors in a single optimization, formulated as MAP estimation in a reduced parameter space
  • Markerless, lighting-free, hardware-light real-time reconstruction of compelling 3D facial dynamics, enabling gameplay and social-interaction applications
Context
Builds on example-based facial rigging (Li et al. 2010) and blendshape/animation-prior models, adapting them to noisy consumer depth-camera input for live retargeting.Builds on: Example-Based Facial Rigging
Correctness
Demonstrated with a commodity 3D sensor; the method assumes pre-recorded animation priors and a reduced expression space to compensate for high sensor noise, so output quality is bounded by the prior and the per-user model rather than capturing arbitrary unseen expressions.
Clarity
Accessible and application-driven; a first pass conveys the live-avatar idea and pipeline, with a second pass for the MAP optimization and the registration/prior terms.
How to read it
First pass for the real-time markerless concept and where the animation priors fit; second pass on the single-optimization formulation if you want the tracking math or to reproduce the stability behavior.

Builds on

Built upon by

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →