Skip to content

← ArchivePaper2024

Democratizing the Creation of Animatable Facial Avatars

Yilin Zhu, Dalton Omens, Haodi He, Ron Fedkiw

arXivAcademic2 citesFacialRigging

Builds a person-specific facial rig without a light stage by warping real photographs onto a template avatar and projecting them into its texture, then capturing expressions with a Simon Says routine.

Abstract

In high-end visual effects pipelines, a customized (and expensive) light stage system is (typically) used to scan an actor in order to acquire both geometry and texture for various expressions. Aiming towards democratization, we propose a novel pipeline for obtaining geometry and texture as well as enough expression information to build a customized person-specific animation rig without using a light stage or any other high-end hardware (or manual cleanup). A key novel idea consists of warping real-world images to align with the geometry of a template avatar and subsequently projecting the warped image into the template avatar's texture; importantly, this allows us to leverage baked-in real-world lighting/texture information in order to create surrogate facial features (and bridge the domain gap) for the sake of geometry reconstruction. Not only can our method be used to obtain a neutral expression geometry and de-lit texture, but it can also be used to improve avatars after they have been imported into an animation system (noting that such imports tend to be lossy, while also hallucinating various features). Since a default animation rig will contain template expressions that do not correctly correspond to those of a particular individual, we use a Simon Says approach to capture various expressions and build a person-specific animation rig (that moves like they do). Our aforementioned warping/projection method has high enough efficacy to reconstruct geometry corresponding to each expressions.

How to read this

Category
Preprint on building a person-specific facial animation rig without a light stage, from Fedkiw's group at Stanford.
Contributions
  • A pipeline that obtains geometry, texture and enough expression information to build a person-specific animation rig with no light stage, no other high-end hardware, and no manual cleanup.
  • The key trick: warp real-world photographs to align with a template avatar's geometry, then project the warped image into that template's texture, which bridges the domain gap by carrying real lighting and texture information along.
  • A de-lit texture and a neutral expression geometry as by-products of that same warping and projection step.
  • A repair use for the same method, improving avatars after an animation system has imported them, since those imports tend to be lossy and to hallucinate features.
  • A Simon Says capture routine that records an individual performing expressions, so the default template expressions are replaced by ones that move the way that person does.
Context
No single predecessor, because the paper argues against a class of pipeline rather than against a specific method. What it is displacing is the light stage, whose lineage in this archive starts at debevec-reflectance-field-2000. The instinct it shares, that ordinary footage plus a good template can stand in for a capture stage, is the one ichim-dynamic-avatar-2015 had already pursued with hand-held video. Read it between those two.Builds on: Acquiring the Reflectance Field of a Human Face · Dynamic 3D Avatar Creation from Hand-Held Video Input
Correctness
A preprint, unreviewed, and its comparison class is pipelines whose output is the industry benchmark, so the honest question is where the quality goes rather than whether it goes. Two places to press. The warping step assumes a template close enough to the subject for the alignment to mean anything, and the paper is thin on how far that stretches. The Simon Says capture depends on the subject performing the requested expressions well, which moves a burden off the hardware and onto the person in front of the camera. Fine wrinkle and pore level detail, the actual reason light stages exist, is the thing to look for in the results and the thing least likely to be there.
Clarity
Readable and argument-forward, with the motivation stated more crisply than the method. The pipeline is described in order and is easy to follow once you get past the parenthetical asides.
How to read it
First pass: abstract and results, then decide whether the quality bar clears your use, because that decides whether the rest is worth your time. Second pass: the warping and projection step, which is the actual contribution, and the Simon Says section, which is where the rig becomes person-specific. Third pass: check whether a reviewed version has appeared since and prefer it, and if you are comparing against a scanned pipeline, put the results next to debevec-reflectance-field-2000 rather than next to a modern neural avatar.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →