Text to Point Cloud Localization with Relation-Enhanced Transformer
Guangzhi Wang, Hehe Fan, Mohan S. Kankanhalli
摘要
Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on a text-to-point-cloud cross-modal localization problem. Given a textual query, it aims to identify the described location from city-scale point clouds. The task involves two challenges. 1) In city-scale point clouds, similar ambient instances may exist in several locations. Searching each location in a huge point cloud with only instances as guidance may lead to less discriminative signals and incorrect results. 2) In textual descriptions, the hints are provided separately. In this case, the relations among those hints are not explicitly described, leaving the difficulties of learning relations to the agent itself. To alleviate the two challenges, we propose a unified Relation-Enhanced Transformer (RET) to improve representation discriminability for both point cloud and nature language queries. The core of the proposed RET is a novel Relation-enhanced Self-Attention (RSA) mechanism, which explicitly encodes instance (hint)-wise relations for the two modalities. Moreover, we propose a fine-grained cross-modal matching method to further refine the location predictions in a subsequent instance-hint matching stage. Experimental results on the KITTI360Pose dataset demonstrate that our approach surpasses the previous state-of-the-art method by large margins.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Text to Point Cloud Localization with Multi-Level Negative Contrastive LearningDunqiang Liu, Shujun Huang, Wen Li, Siqi Shen 等AAAI 2025 · 被引用 7 次
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao 等CVPR 2026 · 被引用 4 次
- Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud LocalizationMingtao Feng, Longlong Mei, Zijie Wu, Jianqiao Luo 等ICCV 2025 · 被引用 2 次
- PointListNet: Deep Learning on 3D Point ListsHehe Fan, Linchao Zhu, Yi Yang, Mohan S. KankanhalliCVPR 2023
- Text2Loc: 3D Point Cloud Localization from Natural LanguageYan Xia, Letian Shi, Zifeng Ding, João F. Henriques 等CVPR 2024
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Rethinking and Improving Relative Position Encoding for Vision TransformerKan Wu, Houwen Peng, Minghao Chen, Jianlong Fu 等ICCV 2021 · 被引用 427 次
相关 Paper
- Text2Pos: Text-to-Point-Cloud Cross-Modal LocalizationManuel Kolmet, Qunjie Zhou, Aljosa Osep, Laura Leal-TaixéCVPR 2022 · 被引用 18 次
- CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based FrameworkYanlong Xu, Haoxuan Qu, Jun Liu, Wenxiao Zhang 等CVPR 2025
- Language Conditioned Spatial Relation Reasoning for 3D Object GroundingShizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid 等NeurIPS 2022 · 被引用 173 次
- 3DVG-Transformer: Relation Modeling for Visual Grounding on Point CloudsLichen Zhao, Daigang Cai, Lu Sheng, Dong XuICCV 2021 · 被引用 234 次
- TransRefer3D: Entity-and-Relation Aware Transformer for Fine-Grained 3D Visual GroundingDailan He, Yusheng Zhao, Junyu Luo, Tianrui Hui 等ACM MM 2021 · 被引用 81 次
