How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
Etai Littwin, Omid Saremi, Madhu Advani, Vimal Thilak, Preetum Nakkiran, Chen Huang, Joshua Susskind
Abstract
Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes using a lightweight predictor network. This is contrasted with the Masked AutoEncoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in the data space rather, than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative). In this work, we seek to understand the mechanism behind this empirical observation by analyzing the training dynamics of deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high-influence features, i.e., features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 518bb23a-dd5c-41b1-ac54-22ce5e92666eCited by top-tier papers7
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero et al.NeurIPS 2025 · 109 citations
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive ArchitecturesHai Huang, Yann LeCun, Randall BalestrieroICLR 2026 · 40 citations
- Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self-Supervised LearningHugues Van Assel, Mark Ibrahim, Tommaso Biancalani, Aviv Regev et al.NeurIPS 2025 · 39 citations
- Semantic Tube Prediction: Beating LLM Data Efficiency with JEPAHai Huang, Yann LeCun, Randall BalestrieroICML 2026 · 8 citations
- Rethinking JEPA: Compute‑Efficient Video Self-Supervised Learning with Frozen TeachersXianhang Li, Chen Huang, Chun-Liang Li, Eran Malach et al.ICLR 2026 · 6 citations
Builds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
Related papers
- Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture – Bridging Predictive and Generative Self-Supervised LearningMoritz Gögl, Christopher YauICML 2026
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular DataHugo Thimonier, José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel et al.ICLR 2025
- Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive ArchitecturesPablo Ruiz-Morales, Dries Vanoost, Davy Pissoort, Mathias VerbekeAAAI 2026
- Text-Conditional JEPA for Learning Semantically Rich Visual RepresentationsChen Huang, Xianhang Li, Vimal Thilak, Etai Littwin et al.ICML 2026 · 1 citation
- VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World ModelsYongchao HuangICML 2026 · 9 citations
