Learning Where to Learn in Cross-View Self-Supervised Learning
Lang Huang, Shan You, Mingkai Zheng, Fei Wang, Chen Qian, Toshihiko Yamasaki
Abstract
Self-supervised learning (SSL) has made enormous progress and largely narrowed the gap with the supervised ones, where the representation learning is mainly guided by a projection into an embedding space. During the projection, current methods simply adopt uniform aggregation of pixels for embedding; however, this risks involving object-irrelevant nuisances and spatial misalignment for different augmentations. In this paper, we present a new approach, Learning Where to Learn (LEWEL), to adaptively aggregate spatial information of features, so that the projected embeddings could be exactly aligned and thus guide the feature learning better. Concretely, we reinterpret the projection head in SSL as a per-pixel projection and predict a set of spatial alignment maps from the original features by this weight-sharing projection head. A spectrum of aligned embeddings is thus obtained by aggregating the features with spatial weighting according to these alignment maps. As a result of this adaptive alignment, we observe substantial improvements on both image-level prediction and dense prediction at the same time: LEWEL improves MoCov2 [15] by 1.6%/1.3%/0.5%/0.4% points, improves BYOL [14] by 1.3%/1.3%/0.7%/0.6% points, on ImageNet linear/semi-supervised classification, Pascal VOC semantic segmentation, and object detection, respectively. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">†</sup> Code: https://t.1y/ZI0A.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf21c612-cdfe-49fc-8fbf-15a9158f2dfcCited by top-tier papers9
- Green Hierarchical Vision Transformer for Masked Image ModelingLang Huang, Shan You, Mingkai Zheng, Fei Wang et al.NeurIPS 2022 · 93 citations
- Towards Unsupervised Domain Generalization for Face Anti-SpoofingYuchen Liu, Yabo Chen, Mengran Gou, Chun-Ting Huang et al.ICCV 2023 · 41 citations
- Multi-Label Self-Supervised Learning with Scene ImagesKe Zhu, Minghao Fu, Jianxin WuICCV 2023 · 21 citations
- SimMatchV2: Semi-Supervised Learning with Graph ConsistencyMingkai Zheng, Shan You, Lang Huang, Chen Luo et al.ICCV 2023 · 17 citations
- Prototype-Based Contrastive Learning with Stage-Wise Progressive Augmentation for Self-Supervised Fine-Grained LearningBaofeng Tan, Xiu-Shen Wei, Lin ZhaoICCV 2025 · 2 citations
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
Related papers
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong et al.CVPR 2021
- Representation Learning by Detecting Incorrect Location EmbeddingsSepehr Sameni, Simon Jenni, Paolo FavaroAAAI 2023 · 8 citations
- Sound and Visual Representation Learning with Multiple Pretraining TasksArun Balajee Vasudevan, Dengxin Dai, Luc Van GoolCVPR 2022 · 4 citations
- Instance Localization for Self-Supervised Detection PretrainingCeyuan Yang, Zhirong Wu, Bolei Zhou, Stephen LinCVPR 2021
- Region Similarity Representation LearningTete Xiao, Colorado J. Reed, Xiaolong Wang, Kurt Keutzer et al.ICCV 2021 · 128 citations
