← ArchivePaper2015
FaceDirector: Continuous Control of Facial Performance in Video
Charles Malleson, Jean-Charles Bazin, Oliver Wang, Derek Bradley, Thabo Beeler, Adrian Hilton, Alexander Sorkine-Hornung
Post-production tool allowing directors to blend between different takes of a facial performance, enabling new emotion transitions via image-space blending.
Abstract
FaceDirector is a method for continuously blending between multiple recorded video takes of an actor performing the same scene with different facial expressions or emotional states, enabling a director to specify arbitrary weighted combinations and smooth transitions in post-production. The approach contributes a robust nonlinear audio-visual synchronization technique that combines normalized facial landmarks and MFCC audio cues in a graph-based cost matrix, removing self-similar ambiguous regions through local cost matrix collapsing to obtain dense frame correspondences between takes. A seamless spatio-temporal blending stage then interpolates timing, facial expression, and local appearance using optical flow warping guided by landmark priors and mask-based compositing back into one source video. The method operates entirely in 2D image space without 3D facial reconstruction, and is demonstrated on emotion transition, performance correction, and timing control applications.
How to read this
- Category
- Method: video-based facial performance blending tool
- Contributions
- FaceDirector, a method to continuously blend between multiple recorded takes of an actor performing the same scene, with director-specified weights and smooth transitions in post-production
- A robust nonlinear audio-visual synchronization technique combining normalized facial landmarks and MFCC audio cues in a graph-based cost matrix, with local cost-matrix collapsing to remove self-similar ambiguities
- A seamless spatio-temporal blending stage using landmark-guided optical-flow warping and mask-based compositing, working entirely in 2D image space without 3D reconstruction
- Context
- Relates to image-based facial editing and performance retiming, deliberately avoiding 3D face reconstruction by operating purely in 2D with audio-visual alignment.
- Correctness
- Demonstrated on emotion transition, performance correction, and timing control; because it is 2D optical-flow and compositing based, results depend on take similarity and landmark/flow accuracy and can break under large pose or appearance differences between takes.
- Clarity
- Approachable; a first pass conveys the align-then-blend idea, a second clarifies the synchronization cost matrix.
- How to read it
- Focus on the audio-visual synchronization (landmarks plus MFCC, cost-matrix collapsing) and the warping/compositing pipeline; second pass on synchronization if you care about robustness across takes.
Related work
- FaceLab: Scalable Facial Performance Capture for Visual Effects 2020 / DigiPro
- 3D Shape Regression for Real-Time Facial Animation 2013 / SIGGRAPH
- It's a UVN Face Rig, Charlie Brown: Facial Techniques for Peanuts 2015 / SIGGRAPH
- Rigid Stabilization of Facial Expressions 2011 / SIGGRAPH
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →