How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
Etai Littwin, Omid Saremi, Madhu Advani, Vimal Thilak, Preetum Nakkiran, Chen Huang, Joshua Susskind
摘要
Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are predictive of each other. A recent successful approach that falls under the JEPA framework is self-distillation, where an online encoder is trained to predict the output of the target encoder, sometimes using a lightweight predictor network. This is contrasted with the Masked AutoEncoder (MAE) paradigm, where an encoder and decoder are trained to reconstruct missing parts of the input in the data space rather, than its latent representation. A common motivation for using the JEPA approach over MAE is that the JEPA objective prioritizes abstract features over fine-grained pixel information (which can be unpredictable and uninformative). In this work, we seek to understand the mechanism behind this empirical observation by analyzing the training dynamics of deep linear models. We uncover a surprising mechanism: in a simplified linear setting where both approaches learn similar representations, JEPAs are biased to learn high-influence features, i.e., features characterized by having high regression coefficients. Our results point to a distinct implicit bias of predicting in latent space that may shed light on its success in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics ModelsUladzislau Sobal, Wancong Zhang, Kyunghyun Cho, Randall Balestriero 等NeurIPS 2025 · 被引用 109 次
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive ArchitecturesHai Huang, Yann LeCun, Randall BalestrieroICLR 2026 · 被引用 40 次
- Joint-Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self-Supervised LearningHugues Van Assel, Mark Ibrahim, Tommaso Biancalani, Aviv Regev 等NeurIPS 2025 · 被引用 39 次
- Semantic Tube Prediction: Beating LLM Data Efficiency with JEPAHai Huang, Yann LeCun, Randall BalestrieroICML 2026 · 被引用 8 次
- Rethinking JEPA: Compute‑Efficient Video Self-Supervised Learning with Frozen TeachersXianhang Li, Chen Huang, Chun-Liang Li, Eran Malach 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
相关 Paper
- Var-JEPA: A Variational Formulation of the Joint-Embedding Predictive Architecture – Bridging Predictive and Generative Self-Supervised LearningMoritz Gögl, Christopher YauICML 2026
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular DataHugo Thimonier, José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel 等ICLR 2025
- Koopman Invariants as Drivers of Emergent Time-Series Clustering in Joint-Embedding Predictive ArchitecturesPablo Ruiz-Morales, Dries Vanoost, Davy Pissoort, Mathias VerbekeAAAI 2026
- Text-Conditional JEPA for Learning Semantically Rich Visual RepresentationsChen Huang, Xianhang Li, Vimal Thilak, Etai Littwin 等ICML 2026 · 被引用 1 次
- VJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World ModelsYongchao HuangICML 2026 · 被引用 9 次
