3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection
Junyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao, Haibing Ren, Hao Shen, Huaxia Xia, Si Liu
Abstract
3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal matching, which is limited by the isolated architecture. In such a paradigm, the detector needs to sample keypoints from raw point clouds due to the inherent properties of 3D point clouds (irregular and large-scale), to generate the corresponding object proposal for each keypoint. However, sparse proposals may leave out the target in detection, while dense proposals may confuse the matching model. Moreover, the language-irrelevant detection stage can only sample a small proportion of keypoints on the target, deteriorating the target prediction. In this paper, we propose a 3D Single-Stage Referred Point Progressive Selection (3D-SPS) method, which progressively selects keypoints with the guidance of language and directly locates the target. Specifically, we propose a Description-aware Keypoint Sampling (DKS) module to coarsely focus on the points of language-relevant objects, which are significant clues for grounding. Besides, we devise a Target-oriented Progressive Mining (TPM) module to finely concentrate on the points of the target, which is enabled by progressive intra-modal relation modeling and inter-modal target mining. 3D-SPS bridges the gap between detection and matching in the 3D visual grounding task, localizing the target at a single stage. Experiments demonstrate that 3D-SPS achieves state-of-the-art performance on both ScanRe-fer and Nr3D/Sr3D datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 877fd054-1076-42ea-b683-214a8315e9e8Cited by top-tier papers62
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng et al.ICCV 2023 · 247 citations
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersHaifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang et al.NeurIPS 2024 · 230 citations
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 157 citations
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang et al.ICLR 2026 · 121 citations
- ViewRefer: Grasp the Multi-view Knowledge for 3D Visual GroundingZoey Guo, Yiwen Tang, Ray Zhang, Dong Wang et al.ICCV 2023 · 86 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
Related papers
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang et al.ICCV 2021 · 115 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.EMNLP 2023 · 7 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.ICCV 2023 · 30 citations
