Skip to content

← ArchivePaper2015

FaceDirector: Continuous Control of Facial Performance in Video

Charles Malleson, Jean-Charles Bazin, Oliver Wang, Derek Bradley, Thabo Beeler, Adrian Hilton, Alexander Sorkine-Hornung

CVPRDisney Research18 citesFacial

Post-production tool allowing directors to blend between different takes of a facial performance, enabling new emotion transitions via image-space blending.

Abstract

FaceDirector is a method for continuously blending between multiple recorded video takes of an actor performing the same scene with different facial expressions or emotional states, enabling a director to specify arbitrary weighted combinations and smooth transitions in post-production. The approach contributes a robust nonlinear audio-visual synchronization technique that combines normalized facial landmarks and MFCC audio cues in a graph-based cost matrix, removing self-similar ambiguous regions through local cost matrix collapsing to obtain dense frame correspondences between takes. A seamless spatio-temporal blending stage then interpolates timing, facial expression, and local appearance using optical flow warping guided by landmark priors and mask-based compositing back into one source video. The method operates entirely in 2D image space without 3D facial reconstruction, and is demonstrated on emotion transition, performance correction, and timing control applications.

How to read this

Category
Method: video-based facial performance blending tool
Contributions
  • FaceDirector, a method to continuously blend between multiple recorded takes of an actor performing the same scene, with director-specified weights and smooth transitions in post-production
  • A robust nonlinear audio-visual synchronization technique combining normalized facial landmarks and MFCC audio cues in a graph-based cost matrix, with local cost-matrix collapsing to remove self-similar ambiguities
  • A seamless spatio-temporal blending stage using landmark-guided optical-flow warping and mask-based compositing, working entirely in 2D image space without 3D reconstruction
Context
Relates to image-based facial editing and performance retiming, deliberately avoiding 3D face reconstruction by operating purely in 2D with audio-visual alignment.
Correctness
Demonstrated on emotion transition, performance correction, and timing control; because it is 2D optical-flow and compositing based, results depend on take similarity and landmark/flow accuracy and can break under large pose or appearance differences between takes.
Clarity
Approachable; a first pass conveys the align-then-blend idea, a second clarifies the synchronization cost matrix.
How to read it
Focus on the audio-visual synchronization (landmarks plus MFCC, cost-matrix collapsing) and the warping/compositing pipeline; second pass on synchronization if you care about robustness across takes.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →