Single Image 3D Object Estimation with Primitive Graph Networks
Qian He, Desen Zhou, Bo Wan, Xuming He
摘要
Reconstructing 3D object from a single image (RGB or depth) is a fundamental problem in visual scene understanding and yet remains challenging due to its ill-posed nature and complexity in real-world scenes. To address those challenges, we adopt a primitive-based representation for 3D object, and propose a two-stage graph network for primitive-based 3D object estimation, which consists of a sequential proposal module and a graph reasoning module. Given a 2D image, our proposal module first generates a sequence of 3D primitives from input image with local feature attention. Then the graph reasoning module performs joint reasoning on a primitive graph to capture the global shape context for each primitive. Such a framework is capable of taking into account rich geometry and semantic constraints during 3D structure recovery, producing 3D objects with more coherent structure even under challenging viewing conditions. We train the entire graph neural network in a stage-wise strategy and evaluate it on three benchmarks: Pix3D, ModelNet and NYU Depth V2. Extensive experiments show that our approach outperforms the previous state of the arts with a considerable margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Coherent 3D Scene Diffusion From a Single RGB ImageManuel Dahnert, Angela Dai, Norman Müller, Matthias NießnerNeurIPS 2024 · 被引用 10 次
- Primitive-Based 3D Human-Object Interaction Modelling and ProgrammingSiqi Liu, Yong-Lu Li, Zhou Fang, Xinpeng Liu 等AAAI 2024 · 被引用 8 次
它引用的顶会 Paper12
- Learning Shape Templates With Structured Implicit FunctionsKyle Genova, Forrester Cole, Daniel Vlasic, Aaron Sarna 等ICCV 2019 · 被引用 427 次
- HOT-Net: Non-Autoregressive Transformer for 3D Hand-Object Pose EstimationLin Huang, Jianchao Tan, Jingjing Meng, Ji Liu 等ACM MM 2020 · 被引用 59 次
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to SentencesZhizhong Han, Chao Chen, Yu-Shen Liu, Matthias ZwickerACM MM 2020 · 被引用 38 次
- LGNN: A Context-aware Line Segment DetectorQuan Meng, Jiakai Zhang, Qiang Hu, Xuming He 等ACM MM 2020 · 被引用 29 次
- MM-Hand: 3D-Aware Multi-Modal Guided Hand Generation for 3D Hand Pose SynthesisZhenyu Wu, Duc Hoang, Shih-Yao Lin, Yusheng Xie 等ACM MM 2020 · 被引用 16 次
相关 Paper
- Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationCheng Zhang, Zhaopeng Cui, Yinda Zhang, Bing Zeng 等CVPR 2021
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung 等CVPR 2020
- Learning Unsupervised Hierarchical Part Decomposition of 3D Objects From a Single RGB ImageDespoina Paschalidou, Luc Van Gool, Andreas GeigerCVPR 2020
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou 等ICCV 2019 · 被引用 373 次
- 3DIAS: 3D Shape Reconstruction with Implicit Algebraic SurfacesMohsen Yavartanoo, Jaeyoung Chung, Reyhaneh Neshatavar, Kyoung Mu LeeICCV 2021 · 被引用 19 次
