Skip to content

← ArchivePaper2023

NeRSemble: Multi-view Radiance Field Reconstruction of Human Heads

Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, Matthias Niessner

SIGGRAPHAcademic238 citesFacial

Dynamic neural radiance fields via hash ensemble deformation for high-fidelity head reconstruction, paired with a 4,700+ sequence multi-view capture dataset.

Abstract

We focus on reconstructing high-fidelity radiance fields of human heads, capturing their animations over time, and synthesizing re-renderings from novel viewpoints at arbitrary time steps. To this end, we propose a new multi-view capture setup composed of 16 calibrated machine vision cameras that record time-synchronized images at 7.1 MP resolution and 73 frames per second. With our setup, we collect a new dataset of over 4700 high-resolution, high-framerate sequences of more than 220 human heads, from which we introduce a new human head reconstruction benchmark1. The recorded sequences cover a wide range of facial dynamics, including head motions, natural expressions, emotions, and spoken language. In order to reconstruct high-fidelity human heads, we propose Dynamic Neural Radiance Fields using Hash Ensembles (NeRSemble). We represent scene dynamics by combining a deformation field and an ensemble of 3D multi-resolution hash encodings. The deformation field allows for precise modeling of simple scene movements, while the ensemble of hash encodings helps to represent complex dynamics. As a result, we obtain radiance field representations of human heads that capture motion over time and facilitate re-rendering of arbitrary novel viewpoints.

How to read this

Category
Method plus dataset/benchmark: dynamic neural radiance fields for human heads
Contributions
  • A multi-view capture rig (16 calibrated cameras, 7.1 MP, 73 fps) and a dataset of 4,700+ sequences across 220+ subjects, released as a head reconstruction benchmark
  • NeRSemble, a dynamic NeRF that pairs a deformation field with an ensemble of multi-resolution hash encodings to model both simple motion and complex dynamics
  • Novel-view, novel-time re-rendering of high-fidelity animated heads
Context
Sits in the neural avatar lineage building on monocular approaches such as Neural Head Avatars (Grassal 2022), pushing toward dense multi-view radiance-field capture.Builds on: Neural Head Avatars from Monocular RGB Videos
Correctness
Validated on the authors' own controlled multi-view capture; results depend on a 16-camera calibrated studio, so quality and applicability to casual or monocular inputs are not claimed.
Clarity
Accessible at a high level; a first pass conveys the capture-plus-hash-ensemble idea, a second pass is needed for the deformation and hash-encoding formulation.
How to read it
First pass for the dataset/benchmark value and the hash-ensemble concept; second pass on the deformation-field plus hash-encoding combination if you intend to reproduce or build on the representation.

Builds on

Built upon by

Nothing yet.

Related work

Keywords

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →