← ArchivePaper2022
EMOCA: Emotion Driven Monocular Face Capture and Animation
Introduces a deep perceptual emotion consistency loss during monocular 3D face reconstruction training, significantly improving the fidelity of reconstructed facial expressions over prior methods.
Abstract
As 3D facial avatars become more widely used for communication, it is critical that they faithfully convey emotion. Unfortunately, the best recent methods that regress parametric 3D face models from monocular images are unable to capture the full spectrum of facial expression, such as subtle or extreme emotions. We find the standard reconstruction metrics used for training (landmark reprojection error, photometric error, and face recognition loss) are insufficient to capture high-fidelity expressions. The result is facial geometries that do not match the emotional content of the input image. We address this with EMOCA (EMOtion Capture and Animation), by introducing a novel deep perceptual emotion consistency loss during training, which helps ensure that the reconstructed 3D expression matches the expression depicted in the input image. While EMOCA achieves 3D reconstruction errors that are on par with the current best methods, it significantly outperforms them in terms of the quality of the reconstructed expression and the perceived emotional content. We also directly regress levels of valence and arousal and classify basic expressions from the estimated 3D face parameters. On the task of in-the-wild emotion recognition, our purely geometric approach is on par with the best image-based methods, highlighting the value of 3D geometry in analyzing human behavior.
How to read this
- Category
- Method: monocular 3D face reconstruction with an emotion-aware training loss
- Contributions
- Introduces a deep perceptual emotion consistency loss so the reconstructed 3D expression matches the emotion in the input image
- Achieves geometric reconstruction error on par with prior best methods while improving perceived expression fidelity
- Additionally regresses valence and arousal and classifies basic expressions from the estimated 3D face parameters
- Context
- Builds on regression-based parametric face capture, extending the in-the-wild detailed-model line of Feng et al.'s DECA with an emotion-driven supervision signal.Builds on: Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
- Correctness
- The central claim is that standard metrics (landmark, photometric, recognition) under-capture expression and that an emotion-consistency loss helps; this is validated largely through reconstruction error parity plus perceptual/emotion measures, so the gain is in perceived expressiveness rather than lower geometric error, and it depends on the quality of the emotion network used for supervision.
- Clarity
- Accessible; a first pass conveys the motivation and the loss idea, with a second pass for the network and training details.
- How to read it
- First pass for the insight that reconstruction metrics miss emotion and how the perceptual loss fixes it; a second pass pays off mainly if you need the training setup or the valence/arousal regression head.
Built upon by
Nothing yet.
Related work
- SPARK: Self-supervised Personalized Real-time Monocular Face Capture 2024 / SIGGRAPH Asia
- Learning an Animatable Detailed 3D Face Model from In-The-Wild Images 2021 / SIGGRAPH
- I M Avatar: Implicit Morphable Head Avatars from Videos 2022 / CVPR
- Towards Metrical Reconstruction of Human Faces 2022 / Eurographics
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →