Augmentation Component Analysis: Modeling Similarity via the Augmentation Overlaps
Lu Han, Han-Jia Ye, De-Chuan Zhan
Abstract
Self-supervised learning aims to learn a embedding space where semantically similar samples are close. Contrastive learning methods pull views of samples together and push different samples away, which utilizes semantic invariance of augmentation but ignores the relationship between samples. To better exploit the power of augmentation, we observe that semantically similar samples are more likely to have similar augmented views. Therefore, we can take the augmented views as a special description of a sample. In this paper, we model such a description as the augmentation distribution and we call it augmentation feature. The similarity in augmentation feature reflects how much the views of two samples overlap and is related to their semantical similarity. Without computational burdens to explicitly estimate values of the augmentation feature, we propose Augmentation Component Analysis (ACA) with a contrastive-like loss to learn principal components and an on-the-fly projection loss to embed data. ACA equals an efficient dimension reduction by PCA and extracts low-dimensional embeddings, theoretically preserving the similarity of augmentation distribution between samples. Empirical results show our method can achieve competitive results against various traditional contrastive learning methods on different benchmarks. Code available at https://github.com/hanlu-nju/AugCA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95d66142-9a3a-4659-96f7-ffffc35bc702Cited by top-tier papers2
- When Can We Approximate Wide Contrastive Models with Neural Tangent Kernels and Principal Component Analysis?Gautham Govind Anil, Pascal Mattia Esser, Debarghya GhoshdastidarAAAI 2025 · 1 citation
- Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame ProjectionsBerken Utku Demirel, Christian HolzNeurIPS 2025 · 1 citation
Builds on22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Contrastive Learning Can Find An Optimal Basis For Approximately View-Invariant FunctionsDaniel D. Johnson, Ayoub El Hanchi, Chris J. MaddisonICLR 2023 · 1 citation
- Revisiting Contrastive Learning through the Lens of Neighborhood Component Analysis: an Integrated FrameworkChing-Yun Ko, Jeet Mohapatra, Sijia Liu, Pin-Yu Chen et al.ICML 2022 · 15 citations
- An Augmentation-Aware Theory for Self-Supervised Contrastive LearningJingyi Cui, Hongwei Wen, Yisen WangICML 2025
- Intriguing Properties of Contrastive LossesTing Chen, Calvin Luo, Lala LiNeurIPS 2021 · 206 citations
- Identity-Disentangled Adversarial Augmentation for Self-supervised LearningKaiwen Yang, Tianyi Zhou, Xinmei Tian, Dacheng TaoICML 2022 · 13 citations
