Skip to content

← ArchivePaper2020

Neural Voice Puppetry: Audio-driven Facial Reenactment

Justus Thies, Mohamed Elgharib, Ayush Tewari, Christian Theobalt, Matthias Niessner

EurographicsAcademic455 cites1 descendantFacial

Audio-driven facial video synthesis via a latent 3D face model space, enabling video dubbing and cross-person reenactment with temporal stability.

Abstract

We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis. Given an audio sequence of a source person or digital assistant, we generate a photo-realistic output video of a target person that is in sync with the audio of the source input. This audio-driven facial reenactment is driven by a deep neural network that employs a latent 3D face model space. Through the underlying 3D representation, the model inherently learns temporal stability while we leverage neural rendering to generate photo-realistic output frames. Our approach generalizes across different people, allowing us to synthesize videos of a target actor with the voice of any unknown source actor or even synthetic voices that can be generated utilizing standard text-to-speech approaches. Neural Voice Puppetry has a variety of use-cases, including audio-driven video avatars, video dubbing, and text-driven video synthesis of a talking head. We demonstrate the capabilities of our method in a series of audio- and text-based puppetry examples, including comparisons to state-of-the-art techniques and a user study.

How to read this

Category
Method: audio-driven facial reenactment
Contributions
  • Audio-driven facial video synthesis that lip-syncs a target person to arbitrary source audio
  • A deep network operating in a latent 3D face model space, with neural rendering for photo-realistic frames
  • Generalizes across people and supports video dubbing, audio-driven avatars, and text-driven talking heads
Context
Extends the audio-to-face line of Karras et al. (joint end-to-end learning of pose and emotion) by routing prediction through a latent 3D face representation and neural rendering rather than rendering directly.Builds on: Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion
Correctness
The 3D intermediate is credited with inherent temporal stability, and claims are backed by comparisons to state-of-the-art methods and a user study; as audio-driven reenactment it carries the usual identity/expression-transfer and potential-misuse caveats, and quality still depends on target-actor footage.
Clarity
Accessible overview; the latent 3D space plus neural rendering pipeline benefits from a second pass.
How to read it
First pass for the use-cases and the role of the latent 3D model; second pass on the audio-to-expression mapping and neural renderer, and skim the user study for perceived quality.

Builds on

Built upon by

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →