Skip to content

← ArchivePaper2003

Vision-based control of 3D facial animation

Jinxiang Chai, Jing Xiao, Jessica K. Hodgins

SCAAcademic200 cites1 descendantFacialRetargeting

Drives a 3D face rig from ordinary video, tracking a performer with a single camera instead of markers and mapping the tracked motion onto the character.

Abstract

Controlling and animating the facial expression of a computer-generated 3D character is a difficult problem because the face has many degrees of freedom while most available input devices have few. In this paper, we show that a rich set of lifelike facial actions can be created from a preprocessed motion capture database and that a user can control these actions by acting out the desired motions in front of a video camera. We develop a real-time facial tracking system to extract a small set of animation control parameters from video. Because of the nature of video data, these parameters may be noisy, low-resolution, and contain errors. The system uses the knowledge embedded in motion capture data to translate these low-quality 2D animation control signals into high-quality 3D facial expressions. To adapt the synthesized motion to a new character model, we introduce an efficient expression retargeting technique whose run-time computation is constant independent of the complexity of the character model. We demonstrate the power of this approach through two users who control and animate a wide range of 3D facial expressions of different avatars.

How to read this

Category
Foundational performance-driven facial animation paper, from the Carnegie Mellon graphics group.
Contributions
  • Control of a 3D facial rig from ordinary video, using a single camera rather than a marker-based capture stage.
  • Vision-based tracking of a performer's face as the animation input signal.
  • A mapping from the tracked, noisy, low-dimensional signal onto the higher-dimensional space of the character rig.
  • A demonstration that consumer-grade input can drive expressive facial animation, two decades before that became routine.
Context
One of the origin points for everything that followed in markerless facial performance capture. It has no ancestor in the archive because it largely is the ancestor: the idea that a cheap camera plus a learned prior can stand in for a capture stage runs from here through lau-face-poser-2009 by the same group and on into the modern avatar work. Its 160 or so citations understate its influence, since much of what it started is now assumed rather than cited.
Correctness
Written in 2003, and its limits are the limits of that moment: single subject, controlled lighting, and a tracking front end that today would be replaced outright by a learned landmark detector. What has aged well is the framing of the problem, mapping a sparse noisy signal onto a rig using data as a prior, which is still how the problem is posed. Judge the paper on the framing, not on the tracker.
Clarity
Well written and still readable. The problem statement is clean and the system diagram carries the argument.
How to read it
First pass: abstract, system diagram, results. Second pass: the mapping from tracked signal to rig parameters, which is the part that is still current, and skim the tracking front end, which is not. Third pass: read it next to lau-face-poser-2009 to see the same group swap real-time tracking for interactive editing under the same kind of prior.

Builds on

Nothing in the archive, this is a starting point.

Built upon by

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →