Single Image 3D Object Estimation with Primitive Graph Networks
Qian He, Desen Zhou, Bo Wan, Xuming He
Abstract
Reconstructing 3D object from a single image (RGB or depth) is a fundamental problem in visual scene understanding and yet remains challenging due to its ill-posed nature and complexity in real-world scenes. To address those challenges, we adopt a primitive-based representation for 3D object, and propose a two-stage graph network for primitive-based 3D object estimation, which consists of a sequential proposal module and a graph reasoning module. Given a 2D image, our proposal module first generates a sequence of 3D primitives from input image with local feature attention. Then the graph reasoning module performs joint reasoning on a primitive graph to capture the global shape context for each primitive. Such a framework is capable of taking into account rich geometry and semantic constraints during 3D structure recovery, producing 3D objects with more coherent structure even under challenging viewing conditions. We train the entire graph neural network in a stage-wise strategy and evaluate it on three benchmarks: Pix3D, ModelNet and NYU Depth V2. Extensive experiments show that our approach outperforms the previous state of the arts with a considerable margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0801cdb-108e-40e8-a80b-078bbcb0c91bCited by top-tier papers2
- Coherent 3D Scene Diffusion From a Single RGB ImageManuel Dahnert, Angela Dai, Norman Müller, Matthias NießnerNeurIPS 2024 · 10 citations
- Primitive-Based 3D Human-Object Interaction Modelling and ProgrammingSiqi Liu, Yong-Lu Li, Zhou Fang, Xinpeng Liu et al.AAAI 2024 · 8 citations
Builds on12
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna et al.ICCV 2019 · 427 citations
- HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose EstimationLin Huang, Jianchao Tan, Jingjing Meng, Ji Liu et al.ACM MM 2020 · 59 citations
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to SentencesZhizhong Han, Chao Chen, Yu-Shen Liu, Matthias ZwickerACM MM 2020 · 38 citations
- LGNN: A Context-aware Line Segment DetectorQuan Meng, Jiakai Zhang, Qiang Hu, Xuming He et al.ACM MM 2020 · 29 citations
- MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose SynthesisZhenyu Wu, Duc Hoang, Shih-Yao Lin, Yusheng Xie et al.ACM MM 2020 · 16 citations
Related papers
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng et al.CVPR 2021
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Learning Unsupervised Hierarchical Part Decomposition of 3D Objects From a Single RGB ImageDespoina Paschalidou, Luc Van Gool, Andreas GeigerCVPR 2020
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
- 3DIAS: 3D Shape Reconstruction with Implicit Algebraic SurfacesMohsen Yavartanoo, Jaeyoung Chung, Reyhaneh Neshatavar, Kyoung Mu LeeICCV 2021 · 19 citations
