Skip to content

← ArchivePaper2022

Towards Metrical Reconstruction of Human Faces

Wojciech Zielonka, Timo Bolkart, Justus Thies

EurographicsAcademic238 citesFacial

MICA predicts metric-accurate FLAME face shapes from single images using supervised learning on 2,000+ identities with face recognition features.

Abstract

Face reconstruction and tracking is a building block of numerous applications in AR/VR, human-machine interac-tion, as well as medical applications. Most of these applications rely on a metrically correct prediction of the shape, especially, when the reconstructed subject is put into a metrical context (i.e., when there is a reference object of known size). A metrical reconstruction is also needed for any application that measures distances and dimensions of the subject (e.g., to virtually fit a glasses frame). State-of-the-art methods for face reconstruction from a single image are trained on large 2D image datasets in a self-supervised fashion. However, due to the nature of a perspective projection they are not able to reconstruct the actual face dimensions, and even predicting the average human face outperforms some of these methods in a metrical sense. To learn the actual shape of a face, we argue for a supervised training scheme. Since there exists no large-scale 3D dataset for this task, we annotated and unified small- and medium-scale databases. The resulting unified dataset is still a medium-scale dataset with more than 2k identities and training purely on it would lead to overfitting.

How to read this

Category
Method: metrically accurate single-image face reconstruction
Contributions
  • MICA, which predicts metric-accurate FLAME face shape from a single image
  • A supervised training scheme arguing that self-supervised 2D-trained methods cannot recover true face dimensions under perspective projection
  • A unified, annotated 3D dataset assembled from small and medium databases (over 2k identities), leveraging face-recognition features
Context
Builds on the FLAME face model (Li et al., Learning a Model of Facial Shape and Expression from 4D Scans, 2017) and contrasts itself with self-supervised single-image methods that lack metric scale.Builds on: Learning a Model of Facial Shape and Expression from 4D Scans
Correctness
The central premise is that metric shape requires supervised 3D training because perspective projection makes 2D self-supervision scale-ambiguous; trained on a unified medium-scale dataset of 2k+ identities, so coverage of identities and conditions outside that data is the natural limitation, and the authors themselves note overfitting risk from training purely on it.
Clarity
Accessible; a first pass conveys the metric-accuracy argument and the dataset strategy, a second pass for the network and loss details.
How to read it
Read once for the metric-vs-self-supervised argument and the dataset unification; second pass on how face-recognition features and the regularization counter overfitting if metric face shape matters to your application.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →