Skip to content

← ArchivePaper2022

I M Avatar: Implicit Morphable Head Avatars from Videos

Yufeng Zheng, Victoria Fernandez Abrevaya, Marcel C. Buhler, Xu Chen, Michael J. Black, Otmar Hilliges

CVPRAcademic305 citesFacial

Implicit head avatar learning from monocular video via neural blendshapes and skinning fields in canonical space with end-to-end analytical gradient training.

Abstract

Traditional 3D morphable face models (3DMMs) provide fine-grained control over expression but cannot easily capture geometric and appearance details. Neural volumetric representations approach photorealism but are hard to animate and do not generalize well to unseen expressions. To tackle this problem, we propose IMavatar (Implicit Morphable avatar), a novel method for learning implicit head avatars from monocular videos. Inspired by the fine-grained control mechanisms afforded by conventional 3DMMs, we represent the expression- and pose-related deformations via learned blendshapes and skinning fields. These attributes are pose-independent and can be used to morph the canonical geometry and texture fields given novel expression and pose parameters. We employ ray marching and iterative root-finding to locate the canonical surface intersection for each pixel. A key contribution is our novel analytical gradient formulation that enables end-to-end training of IMavatars from videos. We show quantitatively and qualitatively that our method improves geometry and covers a more complete expression space compared to state-of-the-art methods. Code and data can be found at https://ait.ethz.ch/projects/2022/IMavatar/.

How to read this

Category
Method: implicit head-avatar learning from monocular video
Contributions
  • IMavatar, learning an implicit morphable head avatar from a monocular video
  • Expression and pose deformations represented via learned blendshapes and skinning fields in a pose-independent canonical space
  • A novel analytical gradient formulation (with ray marching and iterative root-finding) enabling end-to-end training
Context
Bridges traditional 3DMM control with neural volumetric representations, building on detailed 3D face modeling (related to Feng et al.'s DECA, 2021).Builds on: Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
Correctness
Demonstrated quantitatively and qualitatively to improve geometry and expression coverage over prior methods per the abstract; readers should note it is trained per-subject from monocular video and that extrapolating to unseen expressions and finding the canonical surface intersection are the hard, assumption-laden parts.
Clarity
Dense (neural fields plus root-finding plus analytical gradients); a first pass conveys the canonical-space blendshape/skinning idea, deeper passes needed for the gradient derivation.
How to read it
Focus first on the canonical representation and why blendshapes plus skinning fields give control; reserve a careful second/third pass for the analytical gradient and root-finding if reimplementing.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →