Skip to content

← ArchivePaper2026

Generalized Audio-Driven Synthesis of Precise Drummer Motion

Alvaro G. Inesta, Mattia Ryffel, Amit H. Bermano, Robert W. Sumner, Martin Guay

SCADisney ResearchMotion Synthesis

Best Paper Award winner, a generative diffusion framework with a dual objective loss that decouples skeletal integrity from drumstick precision to generalize to in the wild audio.

How to read this

Category
research paper, audio driven generative motion synthesis for instrument performance
Contributions
  • Presents a generative diffusion framework for audio driven full body drumming motion, using a dual objective loss that parameterizes body joints by rotation for skeletal integrity while directly supervising drumstick tip positions in Cartesian coordinates for centimeter level precision
  • Curates a 3.5 plus hour professional drumming motion capture dataset at 120 Hz with a domain specific audio augmentation strategy, so the model generalizes to in the wild, non curated recordings rather than only studio audio
  • Introduces two new evaluation metrics for percussive motion quality, a Percussive Alignment Score measuring audio motion onset correspondence, and an Impact Point Deviation score measuring stick tip hitting fidelity
  • Shows the same audio to motion pipeline can be repurposed as a drum transcriber, outperforming specialized transcription baselines even without task specific training
  • Won the Best Paper award at SCA 2026
Context
This paper sits in the broader line of audio and music conditioned motion synthesis work, dance and other instrument performance generation, but it argues drumming is a fundamentally different problem due to noisy percussive audio, extreme end effector precision requirements, and high accelerations that break the temporal smoothness assumptions of conventional motion diffusion. No builds_on entries are listed, and the paper positions itself as the first data centric, audio conditioned rather than MIDI conditioned, approach to drumming motion, comparing directly against two prior MIDI based drumming systems it cites.
Correctness
Backed by quantitative evaluation against the two prior drumming specific approaches, plus human perceptual studies the paper says found the synthesized motion nearly indistinguishable from ground truth, along with ablations of the dual objective loss and the new PAS and IPD metrics. As with any perceptual study, results depend on the specific viewers and stimuli used, and the in the wild generalization claim rests on the paper's own augmentation strategy and evaluation set rather than a fully independent test.
Clarity
Standard, well structured research paper prose, readable for a character tech or animation researcher with basic familiarity with diffusion models, and the domain framing, why drums differ from melodic instruments, is explained clearly enough for readers without prior music motion background.
How to read it
Five minute pass: read the abstract and look at Figure 1 and Figure 2 for the dual objective loss idea, rotations for the body, Cartesian coordinates for stick tips, and the overall pipeline. Second pass: read the introduction's three contributions plus Section 3 on the motion representation and loss formulation, to see exactly how skeletal integrity is decoupled from stick precision. A third full pass is worthwhile for anyone building audio driven or precision end effector motion synthesis, since the PAS and IPD metrics and the data augmentation recipe in Section 3.2 are reusable outside the drumming domain.

This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →