MultiSiam: Self-supervised Multi-instance Siamese Representation Learning for Autonomous Driving
Kai Chen, Lanqing Hong, Hang Xu, Zhenguo Li, Dit-Yan Yeung
Abstract
Autonomous driving has attracted much attention over the years but turns out to be harder than expected, probably due to the difficulty of labeled data collection for model training. Self-supervised learning (SSL), which leverages unlabeled data only for representation learning, might be a promising way to improve model performance. Existing SSL methods, however, usually rely on the single-centric-object guarantee, which may not be applicable for multi-instance datasets such as street scenes. To alleviate this limitation, we raise two issues to solve: (1) how to define positive samples for cross-view consistency and (2) how to measure similarity in multi-instance circumstances. We first adopt an IoU threshold during random cropping to transfer global-inconsistency to local-consistency. Then, we propose two feature alignment methods to enable 2D feature maps for multi-instance similarity measurement. Addition-ally, we adopt intra-image clustering with self-attention for further mining intra-image similarity and translation-invariance. Experiments show that, when pre-trained on Waymo dataset, our method called Multi-instance Siamese Network (MultiSiam) remarkably improves generalization ability and achieves state-of-the-art transfer performance on autonomous driving benchmarks, including Cityscapes and BDD100K, while existing SSL counterparts like MoCo, MoCo-v2, and BYOL show significant performance drop. By pre-training on SODA10M, a large-scale autonomous driving dataset, MultiSiam exceeds the ImageNet pre-trained MoCo-v2, demonstrating the potential of domain-specific pre-training. Code will be available at https://github.com/KaiChen1998/MultiSiam.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- MagicDrive: Street View Generation with Diverse 3D Geometry ControlRuiyuan Gao, Kai Chen, Enze Xie, Lanqing Hong et al.ICLR 2024 · 248 citations
- Sparse Invariant Risk MinimizationXiao Zhou, Yong Lin, Weizhong Zhang, Tong ZhangICML 2022 · 85 citations
- Model Agnostic Sample Reweighting for Out-of-Distribution LearningXiao Zhou, Yong Lin, Renjie Pi, Weizhong Zhang et al.ICML 2022 · 73 citations
- GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data GenerationKai Chen, Enze Xie, Zhe Chen, Yibo Wang et al.ICLR 2024 · 60 citations
- Learning Super-Features for Image RetrievalPhilippe Weinzaepfel, Thomas Lucas, Diane Larlus, Yannis KalantidisICLR 2022 · 56 citations
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
Related papers
- UniVIP: A Unified Framework for Self-Supervised Visual Pre-trainingZhaowen Li, Yousong Zhu, Fan Yang, Wei Li et al.CVPR 2022 · 29 citations
- Siamese DETRZeren Chen, Gengshi Huang, Wei Li, Jianing Teng et al.CVPR 2023
- Effective Adaptation in Multi-Task Co-Training for Unified Autonomous DrivingXiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu et al.NeurIPS 2022 · 53 citations
- Crafting Better Contrastive Views for Siamese Representation LearningXiangyu Peng, Kai Wang, Zheng Zhu, Mang Wang et al.CVPR 2022 · 107 citations
- Region Similarity Representation LearningTete Xiao, Colorado J. Reed, Xiaolong Wang, Kurt Keutzer et al.ICCV 2021 · 128 citations
