Cross-Modal Match for Language Conditioned 3D Object Grounding
Yachao Zhang, Runze Hu, Ronghui Li, Yanyun Qu, Yuan Xie, Xiu Li
Abstract
Language conditioned 3D object grounding aims to find the object within the 3D scene mentioned by natural language descriptions, which mainly depends on the matching between visual and natural language. Considerable improvement in grounding performance is achieved by improving the multimodal fusion mechanism or bridging the gap between detection and matching. However, several mismatches are ignored, i.e., mismatch in local visual representation and global sentence representation, and mismatch in visual space and corresponding label word space. In this paper, we propose crossmodal match for 3D grounding from mitigating these mismatches perspective. Specifically, to match local visual features with the global description sentence, we propose BEV (Bird’s-eye-view) based global information embedding module. It projects multiple object proposal features into the BEV and the relations of different objects are accessed by the visual transformer which can model both positions and features with long-range dependencies. To circumvent the mismatch in feature spaces of different modalities, we propose crossmodal consistency learning. It performs cross-modal consistency constraints to convert the visual feature space into the label word feature space resulting in easier matching. Besides, we introduce label distillation loss and global distillation loss to drive these matches learning in a distillation way. We evaluate our method in mainstream evaluation settings on three datasets, and the results demonstrate the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1282ab4a-e197-4052-8ece-a0dba8c256ebCited by top-tier papers7
- Learning Commonality, Divergence and Variety for Unsupervised Visible-Infrared Person Re-identificationJiangming Shi, Xiangbo Yin, Yachao Zhang, Zhizhong Zhang et al.NeurIPS 2024 · 36 citations
- UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models PriorYao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo et al.NeurIPS 2024 · 15 citations
- BeyondSparse: Facilitating Mamba to Enhance Cross-Domain 3D Semantic Segmentation in Adverse WeatherYao Wu, Mingwei Xing, Yachao Zhang, Fangyong Wang et al.AAAI 2026 · 1 citation
- MaskViM: Domain Generalized Semantic Segmentation with State Space ModelsJiahao Li, Yang Lu, Yuan Xie, Yanyun QuAAAI 2025 · 1 citation
- PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and SegmentationWenbin Tan, Jiawen Lin, Fangyong Wang, Yuan Xie et al.AAAI 2026
Builds on17
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 234 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid et al.NeurIPS 2022 · 173 citations
Related papers
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu et al.CVPR 2022 · 146 citations
- Multi-Attribute Interactions Matter for 3D Visual GroundingCan Xu, Yuehui Han, Rui Xu, Le Hui et al.CVPR 2024 · 5 citations
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu et al.ACM MM 2020 · 35 citations
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang et al.EMNLP 2021 · 63 citations
