Text2Pos: Text-to-Point-Cloud Cross-Modal Localization
Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixé
Abstract
Natural language-based communication with mobile devices and home appliances is becoming increasingly popular and has the potential to become natural for communicating with mobile robots in the future. Towards this goal, we investigate cross-modal text-to-point-cloud localization that will allow us to specify, for example, a vehicle pick-up or goods delivery location. In particular, we propose Text2Pos, a cross-modal localization module that learns to align textual descriptions with localization cues in a coarse-to-fine manner. Given a point cloud of the environment, Text2Pos locates a position that is specified via a natural language-based description of the immediate surroundings. To train Text2Pos and study its performance, we construct KITTI360Pose, the first dataset for this task based on the recently introduced KITTI360 dataset. Our experiments show that we can localize 65% of textual queries within 15m distance to query locations for top-10 retrieved locations. This is a starting point that we hope will spark future developments towards language-based navigation. “Alexa, hand me over my special delivery at the sidewalk in front of the yellow building next to the blue bus stop.”
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Text to Point Cloud Localization with Relation-Enhanced TransformerGuangzhi Wang, Hehe Fan, Mohan S. KankanhalliAAAI 2023 · 27 citations
- Where am I? Cross-View Geo-localization with Natural Language DescriptionsJunyan Ye, Honglin Lin, Leyan Ou, Dairong Chen et al.ICCV 2025 · 11 citations
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen et al.AAAI 2025 · 7 citations
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li et al.NeurIPS 2025 · 7 citations
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao et al.CVPR 2026 · 4 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang et al.ICCV 2021 · 115 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
Related papers
- Text2Loc: 3D Point Cloud Localization from Natural LanguageYan Xia, Letian Shi, Zifeng Ding, João F. Henriques et al.CVPR 2024
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkYanlong Xu, Haoxuan Qu, Jun Liu, Wenxiao Zhang et al.CVPR 2025
- Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationMingtao Feng, Longlong Mei, Zijie Wu, Jianqiao Luo et al.ICCV 2025 · 2 citations
- Talking Points: Describing and Localizing PixelsMatan Rusanovsky, Shimon Malnick, Shai AvidanICLR 2026
- GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationShangshu Yu, Wen Li, Xiaotian Sun, Zhimin Yuan et al.NeurIPS 2025 · 2 citations
