On the Out-of-Distribution Generalization of Self-Supervised Learning
Wenwen Qiang, Jingyao Wang, Zeen Song, Jiangmeng Li, Changwen Zheng
Abstract
In this paper, we focus on the out-of-distribution (OOD) generalization of self-supervised learning (SSL). By analyzing the mini-batch construction during the SSL training phase, we first give one plausible explanation for SSL having OOD generalization. Then, from the perspective of data generation and causal inference, we analyze and conclude that SSL learns spurious correlations during the training process, which leads to a reduction in OOD generalization. To address this issue, we propose a post-intervention distribution (PID) grounded in the Structural Causal Model. PID offers a scenario where the spurious variable and label variable is mutually independent. Besides, we demonstrate that if each mini-batch during SSL training satisfies PID, the resulting SSL model can achieve optimal worst-case OOD performance. This motivates us to develop a batch sampling strategy that enforces PID constraints through the learning of a latent variable model. Through theoretical analysis, we demonstrate the identifiability of the latent variable model and validate the effectiveness of the proposed sampling strategy. Experiments conducted on various downstream OOD tasks demonstrate the effectiveness of the proposed sampling strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image AlignmentYing Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo et al.ICCV 2025 · 13 citations
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
Builds on29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Causal Balancing for Domain GeneralizationXinyi Wang, Michael Saxon, Jiachen Li, Hongyang Zhang et al.ICLR 2023 · 4 citations
- Breaking Correlation Shift via Conditional Invariant RegularizerMingyang Yi, Ruoyu Wang, Jiacheng Sun, Zhenguo Li et al.ICLR 2023
- Counterfactual Maximum Likelihood Estimation for Training Deep NetworksXinyi Wang, Wenhu Chen, Michael Saxon, William Yang WangNeurIPS 2021 · 9 citations
- Improving Generalization of Dynamic Graph Learning via Environment PromptKuo Yang, Zhengyang Zhou, Qihe Huang, Limin Li et al.NeurIPS 2024 · 14 citations
- Diagnosing and Rectifying Fake OOD Invariance: A Restructured Causal ApproachZiliang Chen, Yongsen Zheng, Zhao-Rong Lai, Quanlong Guan et al.AAAI 2024 · 4 citations
