← ArchivePaper2021
MeshTalk: 3D Face Animation from Speech using Cross-Modality Disentanglement
Alexander Richard, Michael Zollhoefer, Yandong Wen, Fernando de la Torre, Yaser Sheikh
Disentangles audio-correlated and audio-uncorrelated facial motion via a categorical latent space and cross-modality loss for full-face animation.
Abstract
This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their scalability. To improve upon existing models, we propose a generic audio-driven facial animation approach that achieves highly realistic motion synthesis results for the entire face. At the core of our approach is a categorical latent space for facial animation that disentangles audio-correlated and audio-uncorrelated information based on a novel cross-modality loss. Our approach ensures highly accurate lip motion, while also synthesizing plausible animation of the parts of the face that are uncorrelated to the audio signal, such as eye blinks and eye brow motion. We demonstrate that our approach outperforms several baselines and obtains state-of-the-art quality both qualitatively and quantitatively. A perceptual user study demonstrates that our approach is deemed more realistic than the current state-of-the-art in over 75% of cases. We recommend watching the supplemental video before reading the paper: https://github.com/facebookresearch/meshtalk
How to read this
- Category
- Method: speech-driven full-face 3D animation
- Contributions
- A generic (not person-specific) audio-driven approach that synthesizes realistic motion for the entire face from speech
- A categorical latent space with a novel cross-modality loss that disentangles audio-correlated from audio-uncorrelated facial motion
- Accurate lip motion plus plausible animation of audio-uncorrelated parts such as eye blinks and brow motion, reported to be preferred over prior state of the art in a perceptual study
- Context
- Advances the speech-to-face-animation lineage, notably VOCA (Cudeiro et al., Capture, Learning, and Synthesis of 3D Speaking Styles), by addressing static upper-face motion and person-specific limitations through cross-modality disentanglement.Builds on: Capture, Learning, and Synthesis of 3D Speaking Styles
- Correctness
- Key assumption is that facial motion cleanly splits into audio-correlated and audio-uncorrelated components captured by a categorical latent; results rest on quantitative comparisons and a perceptual user study, so readers should weigh the subjective-preference evidence alongside the modelling assumption rather than as a hard accuracy guarantee.
- Clarity
- Accessible motivation; a first pass conveys the disentanglement idea, a second pass is needed for the categorical latent space and loss.
- How to read it
- Read pass one for the cross-modality disentanglement framing; do a second pass on the categorical latent and loss design, and watch the supplemental video, before judging the synthesis quality.
Builds on
Built upon by
Nothing yet.
Related work
- FaceFormer: Speech-Driven 3D Facial Animation with Transformers 2022 / CVPR
- Capture, Learning, and Synthesis of 3D Speaking Styles 2019 / CVPR
- SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation 2023 / CVPR
- MoGlow: Probabilistic and Controllable Motion Synthesis Using Normalising Flows 2020 / SIGGRAPH Asia
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →