LIST: Learning Implicitly from Spatial Transformers for Single-View 3D Reconstruction
Mohammad Samiul Arshad, William J. Beksi
Abstract
Accurate reconstruction of both the geometric and topological details of a 3D object from a single 2D image embodies a fundamental challenge in computer vision. Existing explicit/implicit solutions to this problem struggle to recover self-occluded geometry and/or faithfully reconstruct topological shape structures. To resolve this dilemma, we introduce LIST, a novel neural architecture that leverages local and global image features to accurately reconstruct the geometric and topological structure of a 3D object from a single image. We utilize global 2D features to predict a coarse shape of the target object and then use it as a base for higher-resolution reconstruction. By leveraging both local 2D features from the image and 3D features from the coarse prediction, we can predict the signed distance between an arbitrary point and the target surface via an implicit predictor with great accuracy. Furthermore, our model does not require camera estimation or pixel alignment. It provides an uninfluenced reconstruction from the input-view direction. Through qualitative and quantitative analysis, we show the superiority of our model in reconstructing 3D objects from both synthetic and real-world images against the state of the art. Our source code is publicly available to the research community [13].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6e4faea-6939-4468-a13c-13f7e7a2f2c0Cited by top-tier papers3
- MIDGArD: Modular Interpretable Diffusion over Graphs for Articulated DesignsQuentin Leboutet, Nina Wiedemann, Zhipeng Cai, Michael Paulitsch et al.NeurIPS 2024 · 2 citations
- Adaptive 3D Reconstruction via Diffusion Priors and Forward Curvature-Matching Likelihood UpdatesSeunghyeok Shin, Dabin Kim, Hongki LimNeurIPS 2025 · 1 citation
- View Transformation Robustness for Multi-View 3D Object Reconstruction with Reconstruction Error-Guided View SelectionQi Zhang, Zhouhang Luo, Tao Yu, Hui HuangAAAI 2025 · 1 citation
Builds on14
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- 3D Point Cloud Generative Adversarial Network Based on Tree Structured Graph ConvolutionsDong Wook Shu, Sung Woo Park, Junseok KwonICCV 2019 · 337 citations
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang et al.ICCV 2019 · 218 citations
- Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human ReconstructionTong He, John P. Collomosse, Hailin Jin, Stefano SoattoNeurIPS 2020 · 203 citations
- Deep Meta Functionals for Shape RepresentationGidi Littwin, Lior WolfICCV 2019 · 90 citations
Related papers
- D2IM-Net: Learning Detail Disentangled Implicit Fields From Single ImagesManyi Li, Hao ZhangCVPR 2021
- Front2Back: Single View 3D Shape Reconstruction via Front to Back PredictionYuan Yao, Nico Schertler, Enrique Rosales, Helge Rhodin et al.CVPR 2020
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- 3D Scene Reconstruction With Multi-Layer Depth and Epipolar TransformersDaeyun Shin, Zhile Ren, Erik B. Sudderth, Charless C. FowlkesICCV 2019 · 67 citations
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng et al.CVPR 2021
