Matrix Information Theory for Self-Supervised Learning
Yifan Zhang, Zhiquan Tan, Jingqin Yang, Weiran Huang, Yang Yuan
摘要
The maximum entropy encoding framework provides a unified perspective for many non-contrastive learning methods like SimSiam, Barlow Twins, and MEC. Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encoding loss as matrix uniformity loss. Furthermore, Matrix-SSL enhances the maximum entropy encoding method by seamlessly incorporating matrix alignment loss, directly aligning covariance matrices in different branches. Experimental results reveal that Matrix-SSL outperforms state-of-the-art methods on the ImageNet dataset under linear evaluation settings and on MS-COCO for transfer learning tasks. Specifically, when performing transfer learning tasks on MS-COCO, our method outperforms previous SOTA methods such as MoCo v2 and BYOL up to 3.3% with only 400 epochs compared to 800 epochs pre-training. We also try to introduce representation learning into the language modeling regime by fine-tuning a 7B model using matrix cross-entropy loss, with a margin of 3.1% on the GSM8K dataset over the standard cross-entropy loss. Code available at https://github.com/yifanzhang-pro/Matrix-SSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Diff-eRank: A Novel Rank-Based Metric for Evaluating Large Language ModelsLai Wei, Zhiquan Tan, Chenghai Li, Jindong Wang 等NeurIPS 2024 · 被引用 37 次
- Information Flow in Self-Supervised LearningZhiquan Tan, Jingqin Yang, Weiran Huang, Yang Yuan 等ICML 2024 · 被引用 18 次
- Provable Contrastive Continual LearningYichen Wen, Zhiquan Tan, Kaipeng Zheng, Chuanlong Xie 等ICML 2024 · 被引用 13 次
- OTMatch: Improving Semi-Supervised Learning with Optimal TransportZhiquan Tan, Kaipeng Zheng, Weiran HuangICML 2024 · 被引用 10 次
- Unveiling the Dynamics of Information Interplay in Supervised LearningKun Song, Zhiquan Tan, Bochao Zou, Huimin Ma 等ICML 2024 · 被引用 3 次
它引用的顶会 Paper40
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
相关 Paper
- Self-Supervised Learning via Maximum Entropy CodingXin Liu, Zhongdao Wang, Yali Li, Shengjin WangNeurIPS 2022 · 被引用 65 次
- How Mask Matters: Towards Theoretical Understandings of Masked AutoencodersQi Zhang, Yifei Wang, Yisen WangNeurIPS 2022 · 被引用 119 次
- Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID DataXuanyu Chen, Nan Yang, Shuai Wang, Dong YuanICLR 2026 · 被引用 1 次
- MultiSiam: Self-supervised Multi-instance Siamese Representation Learning for Autonomous DrivingKai Chen, Lanqing Hong, Hang Xu, Zhenguo Li 等ICCV 2021 · 被引用 60 次
- Maximizing Incremental Information Entropy for Contrastive LearningJiansong Zhang, Zhuoqin Yang, Xu Wu, Xiaoling Luo 等ICLR 2026
