← ArchivePaper2015
Vdub: Modifying Face Video of Actors for Plausible Visual Alignment to a Dubbed Audio Track
Pablo Garrido, Levi Valgaerts, Hamid Sarmadi, Ingmar Steiner, Kiran Varanasi, Patrick Perez, Christian Theobalt
Monocular face modification system for dubbing, aligning lip motion to dubbed audio by warping the actor's mouth region.
Abstract
In many countries, foreign movies and TV productions are dubbed, i.e., the original voice of an actor is replaced with a translation that is spoken by a dubbing actor in the country's own language. Dubbing is a complex process that requires specific translations and accurately timed recitations such that the new audio at least coarsely adheres to the mouth motion in the video. However, since the sequence of phonemes and visemes in the original and the dubbing language are different, the video‐to‐audio match is never perfect, which is a major source of visual discomfort. In this paper, we propose a system to alter the mouth motion of an actor in a video, so that it matches the new audio track. Our paper builds on high‐quality monocular capture of 3D facial performance, lighting and albedo of the dubbing and target actors, and uses audio analysis in combination with a space‐time retrieval method to synthesize a new photo‐realistically rendered and highly detailed 3D shape model of the mouth region to replace the target performance. We demonstrate plausible visual quality of our results compared to footage that has been professionally dubbed in the traditional way, both qualitatively and through a user study.
How to read this
- Category
- Method: a monocular face-video editing system for visual dubbing
- Contributions
- Alters an actor's mouth motion in a video so it matches a newly dubbed audio track for plausible lip alignment
- Combines high-quality monocular 3D facial-performance capture (shape, lighting, albedo) with audio analysis and a space-time retrieval method to synthesize a new detailed mouth-region model
- Demonstrates photo-realistic rendering of the edited mouth and compares quality against traditionally, professionally dubbed footage
- Context
- Builds on 3D-face modeling lineage rooted in Blanz and Vetter's morphable model (blanz-morphable-1999), applying monocular performance capture plus audio-driven viseme retrieval to the dubbing problem.Builds on: A Morphable Model for the Synthesis of 3D Faces
- Correctness
- Operates from monocular footage and assumes reliable single-view capture of geometry, lighting and albedo; results are shown qualitatively against professional dubs, so a reader should keep in mind that mouth-region synthesis quality depends on capture accuracy and the audio-to-viseme retrieval match.
- Clarity
- Accessible at a high level; a first pass conveys the pipeline idea, a second pass pays off for the capture and space-time retrieval formulation.
- How to read it
- First pass for the overall pipeline (capture, audio analysis, mouth retrieval, compositing); do a second pass on the retrieval and rendering stages if you care how the new mouth is synthesized and blended.
Builds on
Built upon by
Nothing yet.
Related work
- Reconstruction of Personalized 3D Face Rigs from Monocular Video 2016 / SIGGRAPH
- FaceLab: Scalable Facial Performance Capture for Visual Effects 2020 / DigiPro
- Monocular Facial Performance Capture via Deep Expression Matching 2022 / SCA
- Face2Face: Real-Time Face Capture and Reenactment of RGB Videos 2016 / CVPR
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →