Driving-signal aware full-body avatars
Timur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabián Prada, Takaaki Shiratori, Shih-En Wei, Weipeng Xu, Yaser Sheikh, Jason M. Saragih
Abstract
We present a learning-based method for building driving-signal aware full-body avatars. Our model is a conditional variational autoencoder that can be animated with incomplete driving signals, such as human pose and facial keypoints, and produces a high-quality representation of human geometry and view-dependent appearance. The core intuition behind our method is that better drivability and generalization can be achieved by disentangling the driving signals and remaining generative factors, which are not available during animation. To this end, we explicitly account for information deficiency in the driving signal by introducing a latent space that exclusively captures the remaining information, thus enabling the imputation of the missing factors required during full-body animation, while remaining faithful to the driving signal. We also propose a learnable localized compression for the driving signal which promotes better generalization, and helps minimize the influence of global chance-correlations often found in real datasets. For a given driving signal, the resulting variational model produces a compact space of uncertainty for missing factors that allows for an imputation strategy best suited to a particular application. We demonstrate the efficacy of our approach on the challenging problem of full-body animation for virtual telepresence with driving signals acquired from minimal sensors placed in the environment and mounted on a VR-headset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ebe5a45-e111-4bb0-8e5d-12c71f578105Cited by top-tier papers51
- Structured Local Radiance Fields for Human Avatar ModelingZerong Zheng, Han Huang, Tao Yu, Hongwen Zhang et al.CVPR 2022 · 115 citations
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu et al.ICCV 2021 · 112 citations
- When XR and AI Meet - A Scoping Review on Extended Reality and Artificial IntelligenceTeresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz et al.CHI 2023 · 90 citations
- AvatarReX: Real-time Expressive Full-body AvatarsZerong Zheng, Xiaochen Zhao, Hongwen Zhang, Boning Liu et al.SIGGRAPH 2023 · 80 citations
- VirtualCube: An Immersive 3D Video Communication SystemYizhong Zhang, Jiaolong Yang, Zhen Liu, Ruicheng Wang et al.IEEE VR 2022 · 66 citations
Builds on13
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- MeshSDF: Differentiable Iso-Surface ExtractionEdoardo Remelli, Artem Lukoianov, Stephan R. Richter, Benoît Guillard et al.NeurIPS 2020 · 186 citations
- NPMs: Neural Parametric Models for 3D Deformable ShapesPablo R. Palafox, Aljaz Bozic, Justus Thies, Matthias Nießner et al.ICCV 2021 · 129 citations
Related papers
- Full-Body Motion from a Single Head-Mounted Device: Generating SMPL Poses from Partial ObservationsAndrea Dittadi, Sebastian Dziadzio, Darren Cosker, Ben Lundell et al.ICCV 2021 · 78 citations
- Drivable Volumetric Avatars using Texel-Aligned FeaturesEdoardo Remelli, Timur M. Bagautdinov, Shunsuke Saito, Chenglei Wu et al.SIGGRAPH 2022 · 61 citations
- FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsSadegh Aliakbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon et al.CVPR 2022
- NeuWigs: A Neural Dynamic Model for Volumetric Hair Capture and AnimationZiyan Wang, Giljoo Nam, Tuur Stuyck, Stephen Lombardi et al.CVPR 2023
- Auto-CARD: Efficient and Robust Codec Avatar Driving for Real-time Mobile TelepresenceYonggan Fu, Yuecheng Li, Chenghui Li, Jason M. Saragih et al.CVPR 2023
