Skip to content

← ArchivePaper2023

Neural Face Rigging for Animating and Retargeting Facial Meshes in the Wild

Dafei Qin, Jun Saito, Noam Aigerman, Thibault Groueix, Taku Komura

SIGGRAPHAcademic42 citesFacialRiggingRetargeting

End-to-end deep learning method for automatic rigging and retargeting of arbitrary in-the-wild facial meshes without manual blendshape creation.

Abstract

We propose an end-to-end deep-learning approach for automatic rigging and retargeting of 3D models of human faces in the wild. Our approach, called Neural Face Rigging (NFR), holds three key properties: (i) NFR’s expression space maintains human-interpretable editing parameters for artistic controls; (ii) NFR is readily applicable to arbitrary facial meshes with different connectivity and expressions; (iii) NFR can encode and produce fine-grained details of complex expressions performed by arbitrary subjects. To the best of our knowledge, NFR is the first approach to provide realistic and controllable deformations of in-the-wild facial meshes, without the manual creation of blendshapes or correspondence. We design a deformation autoencoder and train it through a multi-dataset training scheme, which benefits from the unique advantages of two data sources: a linear 3DMM with interpretable control parameters as in FACS and 4D captures of real faces with fine-grained details. Through various experiments, we show NFR’s ability to automatically produce realistic and accurate facial deformations across a wide range of existing datasets and noisy facial scans in-the-wild, while providing artist-controlled, editable parameters.

How to read this

Category
Method: deep-learning facial rigging and retargeting
Contributions
  • NFR, an end-to-end approach that auto-rigs and retargets arbitrary in-the-wild facial meshes without manual blendshapes or correspondence
  • A deformation autoencoder with human-interpretable, FACS-like editing parameters that applies across meshes of differing connectivity
  • A multi-dataset training scheme combining a linear 3DMM (interpretable controls) with 4D real-face captures (fine detail)
Context
Relates to single-scan rig generation work such as Li et al. Dynamic Facial Asset and Rig Generation (2020), aiming to remove manual blendshape/correspondence setup.Builds on: Dynamic Facial Asset and Rig Generation from a Single Scan
Correctness
Demonstrated on existing facial datasets and noisy in-the-wild scans; interpretability and generalization rest on the 3DMM-plus-4D training mix, so behavior outside those distributions should be checked.
Clarity
Reasonably accessible; a first pass conveys the rig-free retargeting goal, a second pass clarifies the autoencoder and multi-dataset training design.
How to read it
First pass to grasp the three claimed properties (interpretable controls, arbitrary meshes, fine detail); second pass on the autoencoder architecture and training scheme if rigging/retargeting is your focus.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →