SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentation
Jiale Huang, Shangfei Wang
Abstract
3D Referring Expression Segmentation (3D-RES) aims to segment objects in point clouds according to language descriptions. Unlike common practices in 2D that utilize learnable query embeddings, recent 3D-RES methods typically generate queries directly from 3D points. However, this direct coupling of queries to raw point clouds introduces new challenges: an impractically large number of queries derived from massive point cloud data and a reliance on non-deterministic sampling algorithms. In this paper, we propose a Semantic-based Adaptive Query Network (SAQN), which introduces a novel query strategy for 3D-RES. Instead of generating queries from points, SAQN employs a learnable query vector for each semantic class. This approach drastically reduces the number of queries while maintaining the advantage of avoiding Hungarian matching through implicit class alignment. Additionally, to address potential cross-object ambiguity within semantic classes, we introduce supplementary queries that are adaptively fused with each class query to disambiguate and enrich representations. Comprehensive experiments show that SAQN achieves state-of-the-art performance while reducing the number of queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a35ebab-a696-40a5-98b1-9bf9bb3f68b2Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Superpoint Transformer for 3D Scene Instance SegmentationJiahao Sun, Chunmei Qing, Junpeng Tan, Xiangmin XuAAAI 2023 · 181 citations
Related papers
- Spatial Matters: Position-Guided 3D Referring Expression SegmentationYabing Wang, Zhuotao Tian, Le Wang, Zheng Qin et al.CVPR 2026
- 3D-GRES: Generalized 3D Referring Expression SegmentationChangli Wu, Yihang Liu, Jiayi Ji, Yiwei Ma et al.ACM MM 2024 · 7 citations
- IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma et al.AAAI 2025 · 5 citations
- RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Qi Chen, Jiayi Ji, Haowei Wang et al.NeurIPS 2024 · 16 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
