IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
Qi Chen, Changli Wu, Jiayi Ji, Yiwei Ma, Danni Yang, Xiaoshuai Sun
Abstract
3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature ambiguity and intent ambiguity. Feature ambiguity arises from information loss or distortion during point cloud acquisition due to limitations such as lighting and viewpoint. Intent ambiguity refers to the model's equal treatment of all queries during the decoding process, lacking top-down task-specific guidance. In this paper, we introduce an Image-enhanced Prompt Decoding Network (IPDN), which leverages multi-view images and task-driven information to enhance the model's reasoning capabilities. To address feature ambiguity, we propose the Multi-view Semantic Embedding (MSE) module, which injects multi-view 2D image information into the 3D scene and compensates for potential spatial information loss. To tackle intent ambiguity, we designed a Prompt-Aware Decoder (PAD) that guides the decoding process by deriving task-driven signals from the interaction between the expression and visual features. Comprehensive experiments demonstrate that IPDN outperforms the state-of-the-art by 1.9 and 4.2 points in mIoU metrics on the 3D-RES and 3D-GRES tasks, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2792441-c6ec-4258-ae59-03abb936e8bdCited by top-tier papers6
- 3D-DRES: Detailed 3D Referring Expression SegmentationQi Chen, Changli Wu, Jiayi Ji, Yiwei Ma et al.AAAI 2026 · 1 citation
- Spatial Matters: Position-Guided 3D Referring Expression SegmentationYabing Wang, Zhuotao Tian, Le Wang, Zheng Qin et al.CVPR 2026
- SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression SegmentationJiale Huang, Shangfei WangCVPR 2026
- Exploring Multimodal Prompts For Unsupervised Continuous Anomaly DetectionMingle Zhou, Jiahui Liu, Jin Wan, Gang Li et al.ACM MM 2025
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang et al.CVPR 2026
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
Related papers
- 3D-GRES: Generalized 3D Referring Expression SegmentationChangli Wu, Yihang Liu, Jiayi Ji, Yiwei Ma et al.ACM MM 2024 · 7 citations
- RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression SegmentationChangli Wu, Qi Chen, Jiayi Ji, Haowei Wang et al.NeurIPS 2024 · 16 citations
- MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression SegmentationChangli Wu, Haodong Wang, Jiayi Ji, Yutian Yao et al.CVPR 2026 · 8 citations
- PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and SegmentationWenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie et al.AAAI 2026
- RefMask3D: Language-Guided Transformer for 3D Referring SegmentationShuting He, Henghui DingACM MM 2024 · 12 citations
