Text to Point Cloud Localization with Multi-Level Negative Contrastive Learning
Dunqiang Liu, Shujun Huang, Wen Li, Siqi Shen, Cheng Wang
Abstract
Language-based localization is a crucial task in robotics and computer vision, enabling robots to understand spatial positions through language. Recent methods rely on contrastive learning to establish correspondences between global features of texts and point clouds. However, the inherent ambiguity of textual descriptions makes it difficult to convey geometric information accurately, forcing alignment of them in the feature space may compromise the expressiveness of the point clouds. Unlike previous methods, this paper proposes using language as a filter to distinguish dissimilar locations. To this end, we propose a robust framework of multi-level negative contrastive learning for language-based localization, fully leveraging the descriptive power of language for spatial localization. Our method learns multiple mismatched factors by minimizing the similarity of different locations at different levels, including global-level, instance-level and relationlevel, respectively. Extensive experiments conducted on the KITTI360Pose benchmark demonstrate that our method outperforms better that the state-of-the-art methods. Specifically, we achieve a 56.3% improvement in Top-1 retrieval recall and a 45.9% improvement in 5m localization recall.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13d78db6-3dae-4905-9271-6e053228b927Cited by top-tier papers2
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao et al.CVPR 2026 · 4 citations
- GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationShangshu Yu, Wen Li, Xiaotian Sun, Zhimin Yuan et al.NeurIPS 2025 · 2 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang et al.ICCV 2021 · 115 citations
- BEVPlace: Learning LiDAR-based Place Recognition using Bird's Eye View ImagesLun Luo, Shuhang Zheng, Yixuan Li, Yongzhi Fan et al.ICCV 2023 · 97 citations
- CASSPR: Cross Attention Single Scan Place RecognitionYan Xia, Mariia Gladkova, Rui Wang, Qianyun Li et al.ICCV 2023 · 72 citations
Related papers
- Text2Pos: Text-to-Point-Cloud Cross-Modal LocalizationManuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-TaixéCVPR 2022 · 18 citations
- Text2Loc: 3D Point Cloud Localization from Natural LanguageYan Xia, Letian Shi, Zifeng Ding, João F. Henriques et al.CVPR 2024
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkYanlong Xu, Haoxuan Qu, Jun Liu, Wenxiao Zhang et al.CVPR 2025
- Text to Point Cloud Localization with Relation-Enhanced TransformerGuangzhi Wang, Hehe Fan, Mohan S. KankanhalliAAAI 2023 · 27 citations
- Soft Contrastive Learning for Visual LocalizationJanine Thoma, Danda Pani Paudel, Luc Van GoolNeurIPS 2020 · 41 citations
