Distilling Coarse-to-Fine Semantic Matching Knowledge for Weakly Supervised 3D Visual Grounding
Zehan Wang, Haifeng Huang, Yang Zhao, Linjun Li, Xize Cheng, Yichen Zhu, Aoxiong Yin, Zhou Zhao
Abstract
3D visual grounding involves finding a target object in a 3D scene that corresponds to a given sentence query. Although many approaches have been proposed and achieved impressive performance, they all require dense object-sentence pair annotations in 3D point clouds, which are both time-consuming and expensive. To address the problem that fine-grained annotated data is difficult to obtain, we propose to leverage weakly supervised annotations to learn the 3D visual grounding model, i.e., only coarse scene-sentence correspondences are used to learn object-sentence links. To accomplish this, we design a novel semantic matching model that analyzes the semantic similarity between object proposals and sentences in a coarse-to-fine manner. Specifically, we first extract object proposals and coarsely select the top-K candidates based on feature and class similarity matrices. Next, we reconstruct the masked keywords of the sentence using each candidate one by one, and the reconstructed accuracy finely reflects the semantic similarity of each candidate to the query. Additionally, we distill the coarse-to-fine semantic matching knowledge into a typical two-stage 3D visual grounding model, which reduces inference costs and improves performance by taking full advantage of the well-studied structure of the existing architectures. We conduct extensive experiments on ScanRefer, Nr3D, and Sr3D, which demonstrate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e2d9b7ad-2e8c-4a6e-bce9-1ca931e6d7fdCited by top-tier papers10
- Chat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersHaifeng Huang, Yilun Chen, Zehan Wang, Rongjie Huang et al.NeurIPS 2024 · 230 citations
- Extending Multi-modal Contrastive RepresentationsZiang Zhang, Zehan Wang, Luping Liu, Rongjie Huang et al.NeurIPS 2024 · 28 citations
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo et al.AAAI 2025 · 12 citations
- 3DRP-Net: 3D Relative Position-aware Network for 3D Visual GroundingZehan Wang, Haifeng Huang, Yang Zhao, Linjun Li et al.EMNLP 2023 · 7 citations
- View-on-Graph: Zero-Shot 3D Visual Grounding via Vision-Language Reasoning on Scene GraphsYuanyuan Liu, Haiyang Mei, Dongyang Zhan, Jiayue Zhao et al.AAAI 2026 · 1 citation
Builds on15
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- Group-Free 3D Object Detection via TransformersZe Liu, Zheng Zhang, Yue Cao, Han Hu et al.ICCV 2021 · 368 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
Related papers
- Exploiting Contextual Objects and Relations for 3D Visual GroundingLi Yang, Chunfeng Yuan, Ziqi Zhang, Zhongang Qi et al.NeurIPS 2023 · 33 citations
- SAT: 2D Semantics Assisted Training for 3D Visual GroundingZhengyuan Yang, Songyang Zhang, Liwei Wang, Jiebo LuoICCV 2021 · 166 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- 3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive SelectionJunyu Luo, Jiahui Fu, Xianghao Kong, Chen Gao et al.CVPR 2022 · 72 citations
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
