← ArchivePaper2022
Monocular Facial Performance Capture via Deep Expression Matching
Stephen Bailey, Jeremy Riviere, Morten Mikkelsen, James F. O'Brien
Deep learning pipeline for monocular video facial capture that matches expressions from RGB input to rigged face models at production quality.
Abstract
Facial performance capture is the process of automatically animating a digital face according to a captured performance of an actor. Recent developments in this area have focused on high‐quality results using expensive head‐scanning equipment and camera rigs. These methods produce impressive animations that accurately capture subtle details in an actor's performance. However, these methods are accessible only to content creators with relatively large budgets. Current methods using inexpensive recording equipment generally produce lower quality output that is unsuitable for many applications. In this paper, we present a facial performance capture method that does not require facial scans and instead animates an artist‐created model using standard blendshapes. Furthermore, our method gives artists high‐level control over animations through a workflow similar to existing commercial solutions. Given a recording, our approach matches keyframes of the video with corresponding expressions from an animated library of poses. A Gaussian process model then computes the full animation by interpolating from the set of matched keyframes. Our expression‐matching method computes a low‐dimensional latent code from an image that represents a facial expression while factoring out the facial identity.
How to read this
- Category
- Method: monocular facial performance capture
- Contributions
- A facial capture method that needs no facial scans and animates an artist-created model using standard blendshapes
- Keyframe expression matching from video against an animated library of poses, with Gaussian-process interpolation for the full animation
- High-level artist control through a workflow similar to existing commercial solutions
- Context
- Aims to lower the cost barrier of deep facial capture (related to Laine et al.'s Production-Level Facial Performance Capture Using Deep CNNs, 2017), targeting creators without large head-scanning budgets.Builds on: Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks
- Correctness
- Trades some fidelity for accessibility using inexpensive recording; the assumption is that a pose library plus expression matching plus Gaussian-process interpolation suffices for production-usable output, so readers should expect quality bounded by the library coverage and the low-dimensional matching.
- Clarity
- Readable and pipeline-focused; a first pass conveys the match-and-interpolate idea, a second pass for the expression-matching and GP details.
- How to read it
- Focus on the keyframe-matching plus GP-interpolation pipeline and the cost-vs-quality tradeoff; second pass on the low-dimensional matching if accuracy matters for your use.
Built upon by
Nothing yet.
Related work
- FaceLab: Scalable Facial Performance Capture for Visual Effects 2020 / DigiPro
- Vdub: Modifying Face Video of Actors for Plausible Visual Alignment to a Dubbed Audio Track 2015 / Eurographics
- Position Manipulation Techniques for Facial Animation 2016 / PhD Thesis
- Production-Level Facial Performance Capture Using Deep Convolutional Neural Networks 2017 / SCA
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →