← ArchivePaper2017
Learning a Model of Facial Shape and Expression from 4D Scans
Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, Javier Romero
FLAME model: articulated jaw, neck, and eyeballs with pose-dependent and expression blendshapes trained on 33,000 3D scans.
Abstract
The field of 3D face modeling has a large gap between high-end and low-end methods. At the high end, the best facial animation is indistinguishable from real humans, but this comes at the cost of extensive manual labor. At the low end, face capture from consumer depth sensors relies on 3D face models that are not expressive enough to capture the variability in natural facial shape and expression. We seek a middle ground by learning a facial model from thousands of accurately aligned 3D scans. Our FLAME model (Faces Learned with an Articulated Model and Expressions) is designed to work with existing graphics software and be easy to fit to data. FLAME uses a linear shape space trained from 3800 scans of human heads. FLAME combines this linear shape space with an articulated jaw, neck, and eyeballs, pose-dependent corrective blendshapes, and additional global expression blendshapes. The pose and expression dependent articulations are learned from 4D face sequences in the D3DFACS dataset along with additional 4D sequences. We accurately register a template mesh to the scan sequences and make the D3DFACS registrations available for research purposes. In total the model is trained from over 33, 000 scans. FLAME is low-dimensional but more expressive than the FaceWarehouse model and the Basel Face Model.
How to read this
- Category
- Method / model: a learned parametric face model
- Contributions
- FLAME, a face model combining a linear shape space with an articulated jaw, neck, and eyeballs plus pose-dependent corrective and global expression blendshapes
- Trained from thousands of aligned 3D scans (shape) and 4D sequences (pose and expression), designed to fit data and work in existing graphics software
- Release of registered D3DFACS sequences for research
- Context
- Builds on the morphable-model tradition of Blanz and Vetter's A Morphable Model for the Synthesis of 3D Faces, adding articulation and learned correctives to bridge high-end and low-end face capture.Builds on: A Morphable Model for the Synthesis of 3D Faces
- Correctness
- Learned from a large scan and 4D corpus and built to be easy to fit; readers should remember the shape and expression spaces are linear (with pose-dependent correctives), so extreme or highly stylized faces outside the training distribution may not be captured.
- Clarity
- Clearly written and widely adopted; a first pass conveys the model structure, a second pass clarifies the registration and training pipeline.
- How to read it
- Focus on how shape, pose, and expression are factored and trained; a second pass on registration and fitting pays off because FLAME is a common building block you will likely reuse or compare against.
Builds on
Built upon by
- Capture, Learning, and Synthesis of 3D Speaking Styles 2019
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image 2019
- Learning an Animatable Detailed 3D Face Model from In-The-Wild Images 2021
- Towards Metrical Reconstruction of Human Faces 2022
- GaussianAvatars: Photorealistic Head Avatars with Rigged 3D Gaussians 2024
Related work
- I M Avatar: Implicit Morphable Head Avatars from Videos 2022 / CVPR
- SPARK: Self-supervised Personalized Real-time Monocular Face Capture 2024 / SIGGRAPH Asia
- Expressive Body Capture: 3D Hands, Face, and Body from a Single Image 2019 / CVPR
- STAR: Sparse Trained Articulated Human Body Regressor 2020 / Eurographics
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →