Skip to content

← ArchivePaper2019

Capture, Learning, and Synthesis of 3D Speaking Styles

Daniel Cudeiro, Timo Bolkart, Cassidy Laidlaw, Anurag Ranjan, Michael J. Black

CVPRAcademic471 cites6 descendantsFacialMotion Synthesis

VOCA: speech-driven 3D facial animation system trained on 29 minutes of 4D scans at 60 fps from 12 speakers, generalizing across identities.

Abstract

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we introduce a unique 4D face dataset with about 29 minutes of 4D scans captured at 60 fps and synchronized audio from 12 speakers. We then train a neural network on our dataset that factors identity from facial motion. The learned model, VOCA (Voice Operated Character Animation) takes any speech signal as input, even speech in languages other than English, and realistically animates a wide range of adult faces. Conditioning on subject labels during training allows the model to learn a variety of realistic speaking styles. VOCA also provides animator controls to alter speaking style, identity-dependent facial shape, and pose (i.e. head, jaw, and eyeball rotations) during animation. To our knowledge, VOCA is the only realistic 3D facial animation model that is readily applicable to unseen subjects without retargeting. This makes VOCA suitable for tasks like in-game video, virtual reality avatars, or any scenario in which the speaker, speech, or language is not known in advance. We make the dataset and model available for research purposes at http://voca.is.tue.mpg.de.

How to read this

Category
Method plus dataset: speech-driven 3D facial animation
Contributions
  • Introduces a 4D face dataset of about 29 minutes of scans at 60 fps with synchronized audio from 12 speakers
  • Trains VOCA, a neural model that factors identity from facial motion and animates unseen adult faces from any speech, including non-English
  • Provides animator controls over speaking style, identity-dependent shape, and head, jaw, and eyeball pose
Context
Builds on audio-driven facial animation (Karras et al., 2017) and the FLAME face model from 4D scans (Li et al., 2017), aiming for a model that applies to unseen subjects without retargeting.Builds on: Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion · Learning a Model of Facial Shape and Expression from 4D Scans
Correctness
Trained on 12 speakers, so identity and style coverage is bounded by that set; the cross-language and unseen-subject generalization is the headline claim, and the field still lacks standard metrics, so realism judgments are partly qualitative.
Clarity
Accessible and well-motivated; a first pass conveys the system and controls, with a second pass for the identity-from-motion factorization.
How to read it
Read for the dataset and the identity-versus-motion factoring; a first pass conveys the capability, do a second pass if you care about how the speaking-style conditioning and animator controls are built.

Builds on

Built upon by

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →