← ArchivePaper2026
TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation
Cheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao, Shi-Min Hu
A Graph CVAE compresses heterogeneous kinematic chains into a shared latent motion manifold, letting motion captured from monocular video be retargeted to characters with arbitrary unseen skeletal topologies.
How to read this
- Category
- generative motion prior and video to animation retargeting paper
- Contributions
- Learns a Universal Motion Manifold with a Graph CVAE that compresses arbitrary kinematic chains into a shared, fixed length latent code conditioned on the target rest pose, disentangling motion dynamics from skeletal topology
- Treats video to animation as a conditional flow matching problem, predicting those topology agnostic latent codes from DINOv3 video features, enabling single pass zero shot retargeting to unseen skeletons without test time optimisation
- Introduces Mobjaverse, a dataset mined from Objaverse-XL through a five stage curation pipeline (kinematic validation, motion standardisation, VLM based semantic filtering, manual verification, texture binding), yielding 5,006 unique skeletal topologies and over 2 million frames
- Reports outperforming template based specialist models (SMPL style) on human and quadruped benchmarks while additionally supporting zero shot retargeting to long tail creatures such as hexapods and inanimate rigged objects
- Context
- TopoCap argues against template based motion capture (SMPL, SMAL, MHR) and classical retargeting methods that assume shared skeleton semantics, neither of which scales to the long tail of generated 3D creatures. The two stage design, a Perceiver style Graph CVAE followed by flow matching, extends the broader diffusion and flow matching motion generation lineage (in the spirit of MDM style models) by making the latent representation topology agnostic through the Mobjaverse dataset built specifically for that purpose.
- Correctness
- The claims rest on training and benchmarking against synthetically rendered videos generated from Mobjaverse, plus a stated zero shot test on real internet video. The dataset's semantic filtering step depends on a vision language model (GPT-5.2) as an automated discriminator, which despite a manual verification pass can still admit mislabeled or noisy assets. The introduction and method sections describe the architecture and data pipeline in detail, but the specific outperforms claim needs checking against the paper's own quantitative tables, which were not part of what was read here.
- Clarity
- Dense SIGGRAPH conference paper prose that assumes familiarity with CVAEs, diffusion and flow matching. Useful for an ML literate technical artist, less approachable for a pure animator without that background.
- How to read it
- First pass, read the abstract, Figure 1 and the three listed contributions to get the topology agnostic pitch in plain terms. Second pass, read the Mobjaverse curation section and the method overview to understand exactly what data and architecture the claim depends on. Third pass, jump to the experiments and results tables to verify the outperforms claim against SMPL based and classical retargeting baselines before trusting it at face value.
This page summarises the entry and links to its original source. The archive never hosts or redistributes the publication itself.Show it in the full archive list →