← ArchivePaper2011
Realtime Performance-Based Facial Animation
Real-time performance-based facial animation driven by depth camera input, enabling live facial retargeting at interactive rates.
Abstract
This paper presents a system for performance-based character animation that enables any user to control the facial expressions of a digital avatar in realtime. The user is recorded in a natural environment using a non-intrusive, commercially available 3D sensor. The simplicity of this acquisition device comes at the cost of high noise levels in the acquired data. To effectively map low-quality 2D images and 3D depth maps to realistic facial expressions, we introduce a novel face tracking algorithm that combines geometry and texture registration with pre-recorded animation priors in a single optimization. Formulated as a maximum a posteriori estimation in a reduced parameter space, our method implicitly exploits temporal coherence to stabilize the tracking. We demonstrate that compelling 3D facial dynamics can be reconstructed in realtime without the use of face markers, intrusive lighting, or complex scanning hardware. This makes our system easy to deploy and facilitates a range of new applications, e.g. in digital gameplay or social interactions.
How to read this
- Category
- System / method: real-time performance-based facial animation from a depth sensor
- Contributions
- A system letting any user drive a digital avatar's facial expressions in real time, recorded with a non-intrusive commercial 3D sensor in a natural environment
- A face tracking algorithm combining geometry and texture registration with pre-recorded animation priors in a single optimization, formulated as MAP estimation in a reduced parameter space
- Markerless, lighting-free, hardware-light real-time reconstruction of compelling 3D facial dynamics, enabling gameplay and social-interaction applications
- Context
- Builds on example-based facial rigging (Li et al. 2010) and blendshape/animation-prior models, adapting them to noisy consumer depth-camera input for live retargeting.Builds on: Example-Based Facial Rigging
- Correctness
- Demonstrated with a commodity 3D sensor; the method assumes pre-recorded animation priors and a reduced expression space to compensate for high sensor noise, so output quality is bounded by the prior and the per-user model rather than capturing arbitrary unseen expressions.
- Clarity
- Accessible and application-driven; a first pass conveys the live-avatar idea and pipeline, with a second pass for the MAP optimization and the registration/prior terms.
- How to read it
- First pass for the real-time markerless concept and where the animation priors fit; second pass on the single-optimization formulation if you want the tracking math or to reproduce the stability behavior.
Builds on
Related work
- Online Modeling for Realtime Facial Animation 2013 / SIGGRAPH
- Direct Manipulation Blendshapes 2010 / TOG
- Expression Packing: As-Few-As-Possible Training Expressions for Blendshape Transfer 2020 / CGF
- Performance-Driven Facial Animation 1990 / SIGGRAPH
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →