Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View Images
Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, Shengping Zhang
Abstract
Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input images sequentially. However, when given the same set of input images with different orders, RNN-based approaches are unable to produce consistent reconstruction results. Moreover, due to long-term memory loss, RNNs cannot fully exploit input images to refine reconstruction results. To solve these problems, we propose a novel framework for single-view and multi-view 3D reconstruction, named Pix2Vox. By using a well-designed encoder-decoder, it generates a coarse 3D volume from each input image. Then, a context-aware fusion module is introduced to adaptively select high-quality reconstructions for each part (e.g., table legs) from different coarse 3D volumes to obtain a fused 3D volume. Finally, a refiner further refines the fused 3D volume to generate the final output. Experimental results on the ShapeNet and Pix3D benchmarks indicate that the proposed Pix2Vox outperforms state-of-the-arts by a large margin. Furthermore, the proposed method is 24 times faster than 3D-R2N2 in terms of backward inference time. The experiments on ShapeNet unseen 3D categories have shown the superior generalization abilities of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc324638-8eea-4962-8a4a-eb6c5340528aCited by top-tier papers68
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
- GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisKatja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas GeigerNeurIPS 2020 · 1,001 citations
- One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationMinghua Liu, Chao Xu, Haian Jin, Linghao Chen et al.NeurIPS 2023 · 755 citations
- Neural Radiance Flow for 4D View Synthesis and Video ProcessingYilun Du, Yinan Zhang, Hong-Xing Yu, Joshua B. Tenenbaum et al.ICCV 2021 · 329 citations
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang et al.CVPR 2026 · 280 citations
Related papers
- Multi-view 3D Reconstruction with TransformersDan Wang, Xinrui Cui, Xun Chen, Zhengxia Zou et al.ICCV 2021 · 111 citations
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Single Image 3D Object Estimation with Primitive Graph NetworksQian He, Desen Zhou, Bo Wan, Xuming HeACM MM 2021 · 1 citation
- Mesh R-CNNGeorgia Gkioxari, Justin Johnson, Jitendra MalikICCV 2019
- PVSeRF: Joint Pixel-, Voxel- and Surface-Aligned Radiance Field for Single-Image Novel View SynthesisXianggang Yu, Jiapeng Tang, Yipeng Qin, Chenghong Li et al.ACM MM 2022 · 6 citations
