← ArchivePaper2022
FDLS: A Deep Learning Approach to Production Quality, Controllable, and Retargetable Facial Performances
Wan-Duo Kurt Ma, Muhammad Ghifary, John P. Lewis, Byungkuk Choi, Haegwang Eom
Deep learning facial performance system combining controllability and retargetability for production-quality digital human faces.
Abstract
Visual effects commonly requires both the creation of realistic synthetic humans as well as retargeting actors’ performances to humanoid characters such as aliens and monsters. Achieving the expressive performances demanded in entertainment requires manipulating complex models with hundreds of parameters. Full creative control requires the freedom to make edits at any stage of the production, which prohibits the use of a fully automatic “black box” solution with uninterpretable parameters. On the other hand, producing realistic animation with these sophisticated models is difficult and laborious. This paper describes FDLS (Facial Deep Learning Solver), which is Weta Digital’s solution to these challenges. FDLS adopts a coarse-to-fine and human-in-the-loop strategy, allowing a solved performance to be verified and (if needed) edited at several stages in the solving process. To train FDLS, we first transform the raw motion-captured data into robust graph features. The feature extraction algorithms were devised after carefully observing the artists’ interpretation of the 3d facial landmarks. Secondly, based on the observation that the artists typically finalize the jaw pass animation before proceeding to finer detail, we solve for the jaw motion first and predict fine expressions with region-based networks conditioned on the jaw position.
How to read this
- Category
- Production method: a deep-learning facial performance solver
- Contributions
- FDLS, Weta Digital's deep-learning solver producing production-quality, controllable, and retargetable facial performances
- A coarse-to-fine, human-in-the-loop strategy letting artists verify and edit a solved performance at several stages
- Transforms raw motion-captured data into robust graph features designed around how artists interpret 3D facial landmarks
- Context
- A VFX production system for retargeting actor performances to humanoid characters, building on deep facial-capture work such as Laine et al.'s Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks.Builds on: Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks
- Correctness
- Production-oriented and validated through studio use rather than as a public benchmark; the design explicitly rejects a black-box solver in favor of interpretable, editable stages, so its strength is controllability in a pipeline and generality outside Weta's tools and data is not claimed.
- Clarity
- Accessible and motivated by practical artist needs; a first pass conveys the coarse-to-fine, human-in-the-loop approach.
- How to read it
- First pass for how controllability and retargetability are balanced and where artists intervene; a second pass pays off for the graph-feature extraction and the staged solving pipeline.
Built upon by
Nothing yet.
Related work
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →