TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
Leyuan Xing, huanjia zhang, Dongyu Pan, Hai Wu, Qiming Xia, Kezheng Xiong, Wen Li, Chenglu Wen, Cheng Wang
Abstract
Reliable navigation and decision-making of autonomous vehicles require both accurate localization and object detection. Traditionally, these two tasks are handled separately, leading to redundant computation and limited crosstask knowledge transfer. This paper proposes TACO, the first Task-Aware COntrastive learning framework, which performs joint LiDAR localization and 3D object detection within a single, unified network. TACO leverages contrastive learning to explicitly decouple and align static geographic features for localization and object-centric features for detection. This bidirectional mutual supervision not only enhances localization robustness in dynamic environments by filtering dynamic noise but also boosts detection accuracy via effective spatial context. Additionally, we propose OxfoLD, the first dataset that provides multitraversal LiDAR localization ground truth with rich 3D object annotations, thereby supporting task validation across various times and weather conditions. Experimental results demonstrate that TACO achieves state-of-the-art localization accuracy while maintaining competitive detection performance. The code and dataset will be publicly available at: https://github.com/xmuxly/OxfoLD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cd54d8a-a7ef-47c4-bd6e-81049fa6b8deBuilds on27
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- Point-to-Voxel Knowledge Distillation for LiDAR Semantic SegmentationYuenan Hou, Xinge Zhu, Yuexin Ma, Chen Change Loy et al.CVPR 2022 · 185 citations
- Learning Multi-Scene Absolute Pose Regression with TransformersYoli Shavit, Ron Ferens, Yosi KellerICCV 2021 · 163 citations
- LidarMultiNet: Towards a Unified Multi-Task Network for LiDAR PerceptionDongqiangzi Ye, Zixiang Zhou, Weijia Chen, Yufei Xie et al.AAAI 2023 · 108 citations
- CASSPR: Cross Attention Single Scan Place RecognitionYan Xia, Mariia Gladkova, Rui Wang, Qianyun Li et al.ICCV 2023 · 72 citations
Related papers
- CALICO: Self-Supervised Camera-LiDAR Contrastive Pre-training for BEV PerceptionJiachen Sun, Haizhong Zheng, Qingzhao Zhang, Atul Prakash et al.ICLR 2024 · 15 citations
- CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object DetectionYanan Zhang, Jiaxin Chen, Di HuangCVPR 2022 · 138 citations
- CO3: Cooperative Unsupervised 3D Representation Learning for Autonomous DrivingRunjian Chen, Yao Mu, Runsen Xu, Wenqi Shao et al.ICLR 2023
- From Dataset to Real-world: General 3D Object Detection via Generalized Cross-domain Few-shot LearningShuangzhi Li, Junlong Shen, Lei Ma, Xingyu LiAAAI 2026
- Towards Universal LiDAR-Based 3D Object Detection by Multi-Domain Knowledge TransferGuile Wu, Tongtong Cao, Bingbing Liu, Xingxin Chen et al.ICCV 2023 · 7 citations
