Waxing-and-Waning: a Generic Similarity-based Framework for Efficient Self-Supervised Learning
Sheng Li, Chao Wu, Ao Li, Yanzhi Wang, Xulong Tang, Geng Yuan
Abstract
Deep Neural Networks (DNNs), essential for diverse applications such as visual recognition and eldercare, often require a large amount of labeled data for training, making widespread deployment of DNNs a challenging task. Self-supervised learning (SSL) emerges as a promising approach, which leverages inherent patterns within data through diverse augmentations to train models without explicit labels. However, while SSL has shown notable advancements in accuracy, its high computation costs remain a daunting impediment, particularly for resourceconstrained platforms. To address this problem, we introduce SIMWNW, a similarity-based efficient self-supervised learning framework. By strategically removing less important regions in augmented images and feature maps, SIMWNW not only reduces computation costs but also eliminates irrelevant features that might slow down the learning process, thereby accelerating model convergence. The experimental results show that SIMWNW effectively reduces the amount of computation costs in self-supervised model training without compromising accuracy. Specifically, SIMWNW yields up to 54% and 51% computation savings in training from scratch and transfer learning tasks, respectively.
Published as a conference paper at ICLR 2024 efficiency. So, it is natural to raise a question: Is there a more general and effective method that can significantly improve the training efficiency of SSL?
Considering the training paradigm of the SSL that leverages different data augmentations on two branches with a siamese encoder model used, it results in a unique property of SSL compared to the conventional supervised learning methods (Tian et al., 2020; Chen et al., 2020a). That is, the augmented input images and feature maps on the two branches inherently have a certain degree of similarity. This is a natural opportunity that could be potentially used for computation saving or simplifying. However, it is insufficient to lead us directly to a simple solution. It is not clear whether similar and dissimilar regions of augmented input images and feature maps in the two branches are equivalently crucial for SSL and whether the similarity remains invariant in low-level features and high-level features. And how can we effectively utilize the similarity to improve training efficiency? Motivated by these questions, we make a comprehensive exploration of the impact of similar regions on SSL accuracy. And we explore two types of methods (i.e., reuse and remove) to exploit the similarities for computation-saving. We find that eliminating the computation on similar regions of augmented input images and each layer's activations can significantly reduce the computation and speed up SSL. To mitigate the region shrinking problem caused by convolution layers, we propose a strategy to effectively and efficiently identify and expand high-similarity regions to ensure a decent overall computation saving. This strategy can be considered a waxing-and-waning process. Putting it all together, we propose our SIMWNW, a generic and efficient SSL framework that can significantly reduce training costs and improve the convergence speed of SSL.
Our SIMWNW framework is generic and can be easily applied to different SSL training methods for training cost-saving. We evaluate our framework in both training from scratch and transfer learning tasks and validate the effectiveness and generalizability of SIMWNW. Specifically, in training from scratch tasks, compared to representative SSL works, SIMWNW provides significant computation savings, peaking at 54% and averaging at 40%, and without sacrificing accuracy. In transfer learning tasks, SIMWNW shows a notable reduction in computation costs, peaking at 51% and averaging at 48%, without accuracy loss. We also compare SIMWNW to efficient SSL approaches. In both training from scratch and transfer learning tasks, SIMWNW consistently outperforms SOTA works, achieving an average computational cost reduction of 18% and 14%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eea6e66f-f92b-4ede-9f81-0ddb20a1c2b0Cited by top-tier papers3
- OASIS: Object-Aware Page Management for Multi-GPU SystemsYueqi Wang, Bingyao Li, Mohamed Tarek Ibn Ziad, Lieven Eeckhout et al.HPCA 2025 · 12 citations
- Content-Aware Dynamic Patchification for Efficient Video DiffusionSheng Li, Connelly Barnes, Mamshad Nayeem Rizve, Hongwu Peng et al.CVPR 2026
- Mutual Effort for Efficiency: A Similarity-based Token Pruning for Vision Transformers in Self-Supervised LearningSheng Li, Qitao Tan, Yue Dai, Zhenglun Kong et al.ICLR 2025
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
Related papers
- Harnessing small projectors and multiple views for efficient vision pretrainingArna Ghosh, Kumar Krishna Agrawal, Shagun Sodhani, Adam Oberman et al.NeurIPS 2024 · 5 citations
- CompRess: Self-Supervised Learning by Compressing RepresentationsSoroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed PirsiavashNeurIPS 2020 · 105 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Improving Representation Learning for Histopathologic Images with Cluster ConstraintsWeiyi Wu, Chongyang Gao, Joseph DiPalma, Soroush Vosoughi et al.ICCV 2023 · 13 citations
- Benchmarking Self-Supervised Learning on Diverse Pathology DatasetsMingu Kang, Heon Song, Seonwook Park, Donggeun Yoo et al.CVPR 2023
