CALICO: Self-Supervised Camera-LiDAR Contrastive Pre-training for BEV Perception
Jiachen Sun, Haizhong Zheng, Qingzhao Zhang, Atul Prakash, Zhuoqing Mao, Chaowei Xiao
Abstract
Perception is crucial in the realm of autonomous driving systems, where bird's eye view (BEV)-based architectures have recently reached state-of-the-art performance. The desirability of self-supervised representation learning stems from the expensive and laborious process of annotating 2D and 3D data. Although previous research has investigated pretraining methods for both LiDAR and camera-based 3D object detection, a unified pretraining framework for multimodal BEV perception is missing. In this study, we introduce CALICO, a novel framework that applies contrastive objectives to both LiDAR and camera backbones. Specifically, CALICO incorporates two stages: point-region contrast (PRC) and region-aware distillation (RAD). PRC better balances the region- and scene-level representation learning on the LiDAR modality and offers significant performance improvement compared to existing methods. RAD effectively achieves contrastive distillation on our self-trained teacher model. CALICO's efficacy is substantiated by extensive evaluations on 3D object detection and BEV map segmentation tasks, where it delivers significant performance improvements. Notably, CALICO outperforms the baseline method by 10.5% and 8.6% on NDS and mAP. Moreover, CALICO boosts the robustness of multimodal 3D object detection against adversarial attacks and corruption. Additionally, our framework can be tailored to different backbones and heads, positioning it as a promising approach for multimodal BEV perception.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov et al.ACM MM 2023 · 7 citations
- CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype LearningRunjian Chen, Hang Zhang, Avinash Ravichandran, Hyoungseob Park et al.ICLR 2026 · 1 citation
- RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object DetectionRui Ding, Zhaonian Kuang, Zongwei Zhou, Meng Yang et al.AAAI 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
Related papers
- Image-to-Lidar Self-Supervised Distillation for Autonomous Driving DataCorentin Sautier, Gilles Puy, Spyros Gidaris, Alexandre Boulch et al.CVPR 2022 · 102 citations
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou et al.CVPR 2023
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie et al.ICCV 2023 · 65 citations
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 21 citations
- TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object DetectionLeyuan Xing, huanjia zhang, Dongyu Pan, Hai Wu et al.CVPR 2026
