MESC-3D: Mining Effective Semantic Cues for 3D Reconstruction from a Single Image
Shaoming Li, Qing Cai, Songqi Kong, Runqing Tan, Heng Tong, Shiji Qiu, Yongguo Jiang, Zhi Liu
Abstract
Figure 1. (a) Previous methods simply performed basic operations on the extracted 2D image information and 3D point cloud without establishing a connection between them. (b) Compared to that, we introduced two key designs: First, the Effective Semantic Mining Module, which effectively mines semantic information from the entangled features and enables point cloud to select the information. Second, the 3D Semantic Prior Learning Module, which aims to enable the model to interpret 3D structures as humans do in 3D reconstruction from a single image. (c) The generalization comparsion between the proposed MESC-3D and SOTA methods on base classes with complex backgrounds. (d) MESC-3D's zero-shot on unseen classes. (e) Comparison with state-of-the-art methods on ShapeNet [1] Dataset on Chamfer Distance (y-axis), parameter count (size of the area), and inference time (x-aixs) which show that MESC-3D achieve the best performance with comparable computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 638eafc6-4e46-4169-bf7a-0f8a0414d9f0Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View ImagesHaozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou et al.ICCV 2019 · 373 citations
Related papers
- 3D Shape Reconstruction from 2D Images with Disentangled Attribute FlowXin Wen, Junsheng Zhou, Yu-Shen Liu, Hua Su et al.CVPR 2022 · 47 citations
- From Image Collections to Point Clouds With Self-Supervised Shape and Pose NetworksNavaneet K. L., Ansu Mathew, Shashank Kashyap, Wei-Chih Hung et al.CVPR 2020
- Unsupervised Learning of Intrinsic Structural Representation PointsNenglun Chen, Lingjie Liu, Zhiming Cui, Runnan Chen et al.CVPR 2020
- Denoise and Contrast for Category Agnostic Shape CompletionAntonio Alliegro, Diego Valsesia, Giulia Fracastoro, Enrico Magli et al.CVPR 2021
- Single View Point Cloud Generation via Unified 3D PrototypeYu Lin, Yigong Wang, Yi-Fan Li, Zhuoyi Wang et al.AAAI 2021 · 11 citations
