Geometric View of Soft Decorrelation in Self-Supervised Learning
Yifei Zhang, Hao Zhu, Zixing Song, Yankai Chen, Xinyu Fu, Ziqiao Meng, Piotr Koniusz, Irwin King
Abstract
Contrastive learning, a form of Self-Supervised Learning (SSL), typically consists of an alignment term and a regularization term. The alignment term minimizes the distance between the embeddings of a positive pair, while the regularization term prevents trivial solutions and expresses prior beliefs about the embeddings. As a widely used regularization technique, soft decorrelation has been employed by several non-contrastive SSL methods to avoid trivial solutions. While the decorrelation term is designed to address the issue of dimensional collapse, we find that it fails to achieve this goal theoretically and experimentally. Based on such a finding, we extend the soft decorrelation regularization to minimize the distance between the covariance matrix and an identity matrix. We provide a new perspective on the geometric distance between positive definite matrices to investigate why the soft decorrelation cannot efficiently solve the dimensional collapse. Furthermore, we construct a family of loss functions utilizing the Bregman Matrix Divergence (BMD), with the soft decorrelation representing a specific instance within this family. We prove that a loss function (LogDet) in this family can solve the issue of dimensional collapse. Our novel loss functions based on BMD exhibit superior performance compared to the soft decorrelation and other baseline techniques, as demonstrated by experimental results on graph and image datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0dee7fb0-5cf0-4bb8-9174-a1284a24214fCited by top-tier papers9
- Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer NetworkXiang Fang, Wanlong Fang, Changshuo Wang, Daizong Liu et al.AAAI 2025 · 10 citations
- Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction FollowingYueen Ma, Dafeng Chi, Shiguang Wu, Yuecheng Liu et al.EMNLP 2025 · 9 citations
- Semi-supervised Node Importance Estimation with Informative Distribution Modeling for Uncertainty RegularizationYankai Chen, Taotao Wang, Yixiang Fang, Yunyu XiaoWWW 2025 · 8 citations
- CrossSpectra: Exploiting Cross-Layer Smoothness for Parameter-Efficient Fine-TuningYifei Zhang, Hao Zhu, Junhao Dong, Haoran Shi et al.NeurIPS 2025 · 5 citations
- Understanding and Mitigating Hyperbolic Dimensional Collapse in Graph Contrastive LearningYifei Zhang, Hao Zhu, Menglin Yang, Jiahong Liu et al.KDD 2025 · 4 citations
Related papers
- Zero-CL: Instance and Feature decorrelation for negative-free symmetric contrastive learningShaofeng Zhang, Feng Zhu, Junchi Yan, Rui Zhao et al.ICLR 2022 · 52 citations
- On Feature Decorrelation in Self-Supervised LearningTianyu Hua, Wenxiao Wang, Zihui Xue, Sucheng Ren et al.ICCV 2021 · 237 citations
- Self-Supervised Learning with an Information Maximization CriterionSerdar Ozsoy, Shadi Hamdan, Sercan Ö. Arik, Deniz Yuret et al.NeurIPS 2022 · 54 citations
- GradGCL: Gradient Graph Contrastive LearningRan Li, Shimin Di, Lei Chen, Xiaofang ZhouICDE 2024 · 3 citations
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
