← ArchivePaper2020
Neural Voice Puppetry: Audio-driven Facial Reenactment
Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, Matthias Niessner
Audio-driven facial video synthesis via a latent 3D face model space, enabling video dubbing and cross-person reenactment with temporal stability.
Abstract
We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis. Given an audio sequence of a source person or digital assistant, we generate a photo-realistic output video of a target person that is in sync with the audio of the source input. This audio-driven facial reenactment is driven by a deep neural network that employs a latent 3D face model space. Through the underlying 3D representation, the model inherently learns temporal stability while we leverage neural rendering to generate photo-realistic output frames. Our approach generalizes across different people, allowing us to synthesize videos of a target actor with the voice of any unknown source actor or even synthetic voices that can be generated utilizing standard text-to-speech approaches. Neural Voice Puppetry has a variety of use-cases, including audio-driven video avatars, video dubbing, and text-driven video synthesis of a talking head. We demonstrate the capabilities of our method in a series of audio- and text-based puppetry examples, including comparisons to state-of-the-art techniques and a user study.
How to read this
- Category
- Method: audio-driven facial reenactment
- Contributions
- Audio-driven facial video synthesis that lip-syncs a target person to arbitrary source audio
- A deep network operating in a latent 3D face model space, with neural rendering for photo-realistic frames
- Generalizes across people and supports video dubbing, audio-driven avatars, and text-driven talking heads
- Context
- Extends the audio-to-face line of Karras et al. (joint end-to-end learning of pose and emotion) by routing prediction through a latent 3D face representation and neural rendering rather than rendering directly.Builds on: Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion
- Correctness
- The 3D intermediate is credited with inherent temporal stability, and claims are backed by comparisons to state-of-the-art methods and a user study; as audio-driven reenactment it carries the usual identity/expression-transfer and potential-misuse caveats, and quality still depends on target-actor footage.
- Clarity
- Accessible overview; the latent 3D space plus neural rendering pipeline benefits from a second pass.
- How to read it
- First pass for the use-cases and the role of the latent 3D model; second pass on the audio-to-expression mapping and neural renderer, and skim the user study for perceived quality.
Related work
- Neural Volumes: Learning Dynamic Renderable Volumes from Images 2019 / SIGGRAPH
- Realtime Performance-Based Facial Animation 2011 / SIGGRAPH
- Audiovisual Inputs for Learning Robust, Real-Time Facial Animation with Lip Sync 2023 / MIG
- Displaced Dynamic Expression Regression for Real-Time Facial Tracking and Animation 2014 / SIGGRAPH
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →