Text2Pos: Text-to-Point-Cloud Cross-Modal Localization
Manuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-Taixé
摘要
Natural language-based communication with mobile devices and home appliances is becoming increasingly popular and has the potential to become natural for communicating with mobile robots in the future. Towards this goal, we investigate cross-modal text-to-point-cloud localization that will allow us to specify, for example, a vehicle pick-up or goods delivery location. In particular, we propose Text2Pos, a cross-modal localization module that learns to align textual descriptions with localization cues in a coarse-to-fine manner. Given a point cloud of the environment, Text2Pos locates a position that is specified via a natural language-based description of the immediate surroundings. To train Text2Pos and study its performance, we construct KITTI360Pose, the first dataset for this task based on the recently introduced KITTI360 dataset. Our experiments show that we can localize 65% of textual queries within 15m distance to query locations for top-10 retrieved locations. This is a starting point that we hope will spark future developments towards language-based navigation. “Alexa, hand me over my special delivery at the sidewalk in front of the yellow building next to the blue bus stop.”
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Text to Point Cloud Localization with Relation-Enhanced TransformerGuangzhi Wang, Hehe Fan, Mohan S. KankanhalliAAAI 2023 · 被引用 27 次
- Where am I? Cross-View Geo-localization with Natural Language DescriptionsJunyan Ye, Honglin Lin, Leyan Ou, Dairong Chen 等ICCV 2025 · 被引用 11 次
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen 等AAAI 2025 · 被引用 7 次
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event CamerasLingdong Kong, Dongyue Lu, Alan Liang, Rong Li 等NeurIPS 2025 · 被引用 7 次
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual ReferringZhihao Yuan, Xu Yan, Yinghong Liao, Ruimao Zhang 等ICCV 2021 · 被引用 188 次
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang 等ICCV 2021 · 被引用 115 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
相关 Paper
- Text2Loc: 3D Point Cloud Localization from Natural LanguageYan Xia, Letian Shi, Zifeng Ding, João F. Henriques 等CVPR 2024
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkYanlong Xu, Haoxuan Qu, Jun Liu, Wenxiao Zhang 等CVPR 2025
- Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationMingtao Feng, Longlong Mei, Zijie Wu, Jianqiao Luo 等ICCV 2025 · 被引用 2 次
- Talking Points: Describing and Localizing PixelsMatan Rusanovsky, Shimon Malnick, Shai AvidanICLR 2026
- GTR-Loc: Geospatial Text Regularization Assisted Outdoor LiDAR LocalizationShangshu Yu, Wen Li, Xiaotian Sun, Zhimin Yuan 等NeurIPS 2025 · 被引用 2 次
