ChallenCap: Monocular 3D Capture of Challenging Human Performances Using Multi-Modal References
Yannan He, Anqi Pang, Xin Chen, Han Liang, Minye Wu, Yuexin Ma, Lan Xu
Abstract
Capturing challenging human motions is critical for numerous applications, but it suffers from complex motion patterns and severe self-occlusion under the monocular setting. In this paper, we propose ChallenCap -a template-based approach to capture challenging 3D human motions using a single RGB camera in a novel learning-and-optimization framework, with the aid of multi-modal references. We propose a hybrid motion inference stage with a generation network, which utilizes a temporal encoder-decoder to extract the motion details from the pair-wise sparse-view reference, as well as a motion discriminator to utilize the unpaired marker-based references to extract specific challenging motion characteristics in a data-driven manner. We further adopt a robust motion optimization stage to increase the tracking accuracy, by jointly utilizing the learned motion details from the supervised multi-modal references as well as the reliable motion hints from the input image reference. Extensive experiments on our new challenging motion dataset demonstrate the effectiveness and robustness of our approach to capture challenging human motions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af96402e-6b84-45ac-9b9d-14d7f8935a2fCited by top-tier papers17
- Fourier PlenOctrees for Dynamic Radiance Field Rendering in Real-timeLiao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao et al.CVPR 2022 · 113 citations
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu et al.ICCV 2021 · 112 citations
- Editable free-viewpoint video using a layered neural representationJiakai Zhang, Xinhang Liu, Xinyi Ye, Fuqiang Zhao et al.SIGGRAPH 2021 · 80 citations
- LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR SensorsYiming Ren, Chengfeng Zhao, Yannan He, Peishan Cong et al.IEEE VR 2023 · 50 citations
- LiDARCap: Long-range Markerless 3D Human Motion Capture with LiDAR Point CloudsJialian Li, Jingyi Zhang, Zhiyong Wang, Siqi Shen et al.CVPR 2022 · 49 citations
Builds on16
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
- Tex2Shape: Detailed Full Human Body Geometry From a Single ImageThiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, Marcus A. MagnorICCV 2019 · 343 citations
Related papers
- EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event CameraLan Xu, Weipeng Xu, Vladislav Golyanik, Marc Habermann et al.CVPR 2020
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
- Towards Unstructured Unlabeled Optical Mocap: A Video Helps!Nicholas Milef, John Keyser, Shu KongSIGGRAPH 2024 · 2 citations
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- I'M HOI: Inertia-Aware Monocular Capture of 3D Human-Object InteractionsChengfeng Zhao, Juze Zhang, Jiashen Du, Ziwei Shan et al.CVPR 2024 · 9 citations
