Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self-Supervised Learning
Hugues Van Assel, Mark Ibrahim, Tommaso Biancalani, Aviv Regev, Randall Balestriero
Abstract
Reconstruction and joint-embedding have emerged as two leading paradigms in Self-Supervised Learning (SSL). Reconstruction methods focus on recovering the original sample from a different view in input space. On the other hand, joint-embedding methods align the representations of different views in latent space. Both approaches offer compelling advantages, yet practitioners lack clear guidelines for choosing between them. In this work, we unveil the core mechanisms that distinguish each paradigm. By leveraging closed-form solutions for both approaches, we precisely characterize how the view generation process, e.g. data augmentation, impacts the learned representations. We then demonstrate that, unlike supervised learning, both SSL paradigms require a minimal alignment between augmentations and irrelevant features to achieve asymptotic optimality with increasing sample size. Our findings indicate that in scenarios where these irrelevant features have a large magnitude, joint-embedding methods are preferable because they impose a strictly weaker alignment condition compared to reconstruction-based methods. These results not only clarify the trade-offs between the two paradigms but also substantiate the empirical success of joint-embedding approaches on real-world challenging datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e6130b2-02a2-43b9-a822-5642a54acf0fCited by top-tier papers7
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero et al.ICML 2026 · 19 citations
- Omni-fMRI: A Universal Atlas-Free fMRI Foundation ModelMo Wang, Wenhao Ye, Junfeng Xia, Junxiang Zhang et al.ICML 2026 · 7 citations
- FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training SpeedJiaqi Zhang, Juntuo Wang, Zhixin Sun, John Zou et al.NeurIPS 2025 · 5 citations
- Learning Sparse Visual Representations via Spatial-Semantic FactorizationTheodore Z. Zhao, Sid Kiblawi, Jianwei Yang, Naoto Usuyama et al.ICML 2026
- Reconstruction Outcomes Look Similar but Processes Differ: Improving Context Consistency and Coverage in Graph Masked Auto-EncoderGeng Tang, Keyu Liu, Xibei Yang, Yuhua QianICML 2026
Builds on31
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- Can Generative Models Improve Self-Supervised Representation Learning?Sana Ayromlou, Vahid Reza Khazaie, Fereshteh Forghani, Arash AfkanpourAAAI 2025 · 5 citations
- RankMe: Assessing the Downstream Performance of Pretrained Self-Supervised Representations by Their RankQuentin Garrido, Randall Balestriero, Laurent Najman, Yann LeCunICML 2023 · 127 citations
- Memorization in Self-Supervised Learning Improves Downstream GeneralizationWenhao Wang, Muhammad Ahmad Kaleem, Adam Dziedzic, Michael Backes et al.ICLR 2024 · 19 citations
- You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised LearningThéo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou et al.NeurIPS 2024 · 23 citations
- Analyzing Data-Centric Properties for Graph Contrastive LearningPuja Trivedi, Ekdeep Singh Lubana, Mark Heimann, Danai Koutra et al.NeurIPS 2022 · 13 citations
