← ArchivePaper2022
I M Avatar: Implicit Morphable Head Avatars from Videos
Yufeng Zheng, Victoria Fernandez Abrevaya, Marcel C. Buhler, Xu Chen, Michael J. Black, Otmar Hilliges
Implicit head avatar learning from monocular video via neural blendshapes and skinning fields in canonical space with end-to-end analytical gradient training.
Abstract
Traditional 3D morphable face models (3DMMs) provide fine-grained control over expression but cannot easily capture geometric and appearance details. Neural volumetric representations approach photorealism but are hard to animate and do not generalize well to unseen expressions. To tackle this problem, we propose IMavatar (Implicit Morphable avatar), a novel method for learning implicit head avatars from monocular videos. Inspired by the fine-grained control mechanisms afforded by conventional 3DMMs, we represent the expression- and pose-related deformations via learned blendshapes and skinning fields. These attributes are pose-independent and can be used to morph the canonical geometry and texture fields given novel expression and pose parameters. We employ ray marching and iterative root-finding to locate the canonical surface intersection for each pixel. A key contribution is our novel analytical gradient formulation that enables end-to-end training of IMavatars from videos. We show quantitatively and qualitatively that our method improves geometry and covers a more complete expression space compared to state-of-the-art methods. Code and data can be found at https://ait.ethz.ch/projects/2022/IMavatar/.
How to read this
- Category
- Method: implicit head-avatar learning from monocular video
- Contributions
- IMavatar, learning an implicit morphable head avatar from a monocular video
- Expression and pose deformations represented via learned blendshapes and skinning fields in a pose-independent canonical space
- A novel analytical gradient formulation (with ray marching and iterative root-finding) enabling end-to-end training
- Context
- Bridges traditional 3DMM control with neural volumetric representations, building on detailed 3D face modeling (related to Feng et al.'s DECA, 2021).Builds on: Learning an Animatable Detailed 3D Face Model from In-The-Wild Images
- Correctness
- Demonstrated quantitatively and qualitatively to improve geometry and expression coverage over prior methods per the abstract; readers should note it is trained per-subject from monocular video and that extrapolating to unseen expressions and finding the canonical surface intersection are the hard, assumption-laden parts.
- Clarity
- Dense (neural fields plus root-finding plus analytical gradients); a first pass conveys the canonical-space blendshape/skinning idea, deeper passes needed for the gradient derivation.
- How to read it
- Focus first on the canonical representation and why blendshapes plus skinning fields give control; reserve a careful second/third pass for the analytical gradient and root-finding if reimplementing.
Built upon by
Nothing yet.
Related work
- PointAvatar: Deformable Point-Based Head Avatars from Videos 2023 / CVPR
- SPARK: Self-supervised Personalized Real-time Monocular Face Capture 2024 / SIGGRAPH Asia
- 3D Gaussian Blendshapes for Head Avatar Animation 2024 / SIGGRAPH
- Learning a Model of Facial Shape and Expression from 4D Scans 2017 / SIGGRAPH Asia
Keywords
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →