Text2Loc: 3D Point Cloud Localization from Natural Language
Yan Xia, Letian Shi, Zifeng Ding, João F. Henriques, Daniel Cremers
Abstract
Hi, I am standing on the west of a green building, east of a green road, west of a black garage... Got it! Coming soon! Localization Recall (%) Number of top retrievals Figure 1. (Left) We introduce Text2Loc, a solution designed for city-scale position localization using textual descriptions. When provided with a point cloud representing the surroundings and a textual query describing a position, Text2Loc determines the most probable location of that described position within the map. (Right) Localization performance on the KITTI360Pose test set. The proposed Text2Loc achieves consistently better performance across all top retrieval numbers. Notably, it outperforms the best baseline by up to 2 times, localizing text queries below 5 m.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Where am I? Cross-View Geo-localization with Natural Language DescriptionsJunyan Ye, Honglin Lin, Leyan Ou, Dairong Chen et al.ICCV 2025 · 11 citations
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen et al.AAAI 2025 · 7 citations
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao et al.CVPR 2026 · 4 citations
- L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing ImageryZiwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia et al.NeurIPS 2025 · 3 citations
- GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationShangshu Yu, Wen Li, Xiaotian Sun, Zhimin Yuan et al.NeurIPS 2025 · 2 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang et al.ICCV 2021 · 188 citations
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang et al.ICCV 2021 · 115 citations
- SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place RecognitionZhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu et al.AAAI 2022 · 95 citations
- CASSPR: Cross Attention Single Scan Place RecognitionYan Xia, Mariia Gladkova, Rui Wang, Qianyun Li et al.ICCV 2023 · 72 citations
Related papers
- Text2Pos: Text-to-Point-Cloud Cross-Modal LocalizationManuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-TaixéCVPR 2022 · 18 citations
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkYanlong Xu, Haoxuan Qu, Jun Liu, Wenxiao Zhang et al.CVPR 2025
- Text to Point Cloud Localization with Relation-Enhanced TransformerGuangzhi Wang, Hehe Fan, Mohan S. KankanhalliAAAI 2023 · 27 citations
- Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationMingtao Feng, Longlong Mei, Zijie Wu, Jianqiao Luo et al.ICCV 2025 · 2 citations
- EP2P-Loc: End-to-End 3D Point to 2D Pixel Localization for Large-Scale Visual LocalizationMinjung Kim, Junseo Koo, Gunhee KimICCV 2023 · 22 citations
