← ArchivePaper2023
Audiovisual Inputs for Learning Robust, Real-Time Facial Animation with Lip Sync
Iñaki Navarro, Dario Kneubuehler, Tijmen Verhulsdonck, Eloi du Bois, William Welch, Charles Shang, Ian Sachs, Morgan McGuire, Victor B. Zordan, Kiran Bhat
Combines video and audio streams via two specialized networks to drive facial character rigs in real time on low-end devices.
Abstract
We present an approach for generating facial animation that combines video and audio input data in real time for low-end devices through deep learning. Our method produces control signals from audiovisual inputs separately, and mixes them to animate a character rig. The architecture relies on two specialized networks that are trained on a combination of synthetic and real world data and are highly engineered to be efficient in order to support quality avatar faces even on low-end devices. In addition, the system supports several levels of detail that degrade gracefully for additional scaling and efficiency. We showcase how user testing has been employed to improve performance and a comparison with state of the art.
How to read this
- Category
- Method: a real-time audiovisual facial animation system
- Contributions
- An approach combining video and audio inputs in real time to drive a character rig on low-end devices
- Two specialized, efficiency-engineered networks that produce control signals separately and then mix them
- Training on combined synthetic and real-world data, with graceful level-of-detail degradation for scaling
- Context
- Relates to speech-driven facial synthesis such as Cudeiro's VOCA, adding a video stream and an efficiency focus for real-time, low-end-device deployment.Builds on: Capture, Learning, and Synthesis of 3D Speaking Styles
- Correctness
- Trained on a mix of synthetic and real data and refined through user testing with a state-of-the-art comparison; quality and efficiency claims are tied to the targeted devices and rigs, and the synthetic-data reliance is a caveat for generalization.
- Clarity
- Accessible as a systems/engineering paper; a first pass conveys the dual-network design without heavy math.
- How to read it
- Focus on the two-network split and the audio/video mixing plus the level-of-detail strategy; a second pass is mainly worthwhile for the efficiency engineering if you deploy on constrained devices.
Builds on
Built upon by
Nothing yet.
Related work
- Animating Facial Expressions 1981 / SIGGRAPH
- Sketch-Based Controllers for Blendshape Facial Animation 2015 / Eurographics
- Easy Generation of Facial Animation Using Motion Graphs 2017 / CGF
- Practice and Theory of Blendshape Facial Models 2014 / Eurographics
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →