Skip to content

← ArchivePaper2023

Audiovisual Inputs for Learning Robust, Real-Time Facial Animation with Lip Sync

Iñaki Navarro, Dario Kneubuehler, Tijmen Verhulsdonck, Eloi du Bois, William Welch, Charles Shang, Ian Sachs, Morgan McGuire, Victor B. Zordan, Kiran Bhat

MIGAcademic1 citesFacialMotion Synthesis

Combines video and audio streams via two specialized networks to drive facial character rigs in real time on low-end devices.

Abstract

We present an approach for generating facial animation that combines video and audio input data in real time for low-end devices through deep learning. Our method produces control signals from audiovisual inputs separately, and mixes them to animate a character rig. The architecture relies on two specialized networks that are trained on a combination of synthetic and real world data and are highly engineered to be efficient in order to support quality avatar faces even on low-end devices. In addition, the system supports several levels of detail that degrade gracefully for additional scaling and efficiency. We showcase how user testing has been employed to improve performance and a comparison with state of the art.

How to read this

Category
Method: a real-time audiovisual facial animation system
Contributions
  • An approach combining video and audio inputs in real time to drive a character rig on low-end devices
  • Two specialized, efficiency-engineered networks that produce control signals separately and then mix them
  • Training on combined synthetic and real-world data, with graceful level-of-detail degradation for scaling
Context
Relates to speech-driven facial synthesis such as Cudeiro's VOCA, adding a video stream and an efficiency focus for real-time, low-end-device deployment.Builds on: Capture, Learning, and Synthesis of 3D Speaking Styles
Correctness
Trained on a mix of synthetic and real data and refined through user testing with a state-of-the-art comparison; quality and efficiency claims are tied to the targeted devices and rigs, and the synthetic-data reliance is a caveat for generalization.
Clarity
Accessible as a systems/engineering paper; a first pass conveys the dual-network design without heavy math.
How to read it
Focus on the two-network split and the audio/video mixing plus the level-of-detail strategy; a second pass is mainly worthwhile for the efficiency engineering if you deploy on constrained devices.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →