SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer Fusion
Chuanmao Fan, Chenxi Zhao, Ye Duan
Abstract
Recently, the community has witnessed significant progress in human modeling from single or multi-view inputs. However, these approaches often rely on guessing the occluded regions through either generative models or template fitting. In this work, we address this challenge by exploring optimal fusion strategies using only sparse multi-view inputs. We propose SMVRT, an end-to-end implicit 3D reconstruction framework for sparse multi-view human modeling. Our key contribution lies in the fusion blocks at three stages of the network. First, local and global features alternating global and local fusion modules are designed to enhance 2D features. Second, attentional fusion is performed on warped multi-view and multi-level 2D features to form 3D feature grid. The feature grid aggregates spatially coherent multi-view features by 3D regularization. Third, attentional 2D-3D feature aggregation generates the enhanced latent embeddings for query points to decode occupancies. Experiments on the THUman 2.0/2.1, MultiGarment and Mul-tiHuman datasets demonstrate that our system significantly outperforms state-of-the-art methods both qualitatively and quantitatively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46e8d5bd-3c81-40ee-9344-b2efdf3aedb2Builds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Multi-Garment Net: Learning to Dress 3D People From ImagesBharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, Gerard Pons-MollICCV 2019 · 447 citations
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum et al.NeurIPS 2021 · 426 citations
Related papers
- Learning Visibility Field for Detailed 3D Human Reconstruction and RelightingRuichen Zheng, Peng Li, Haoqian Wang, Tao YuCVPR 2023
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu et al.ICCV 2021 · 112 citations
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao et al.NeurIPS 2025 · 8 citations
- Multi-Person Implicit Reconstruction From a Single ImageArmin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian HiltonCVPR 2021
- CrossHuman: Learning Cross-guidance from Multi-frame Images for Human ReconstructionLiliang Chen, Jiaqi Li, Han Huang, Yandong GuoACM MM 2022 · 4 citations
