You Don't Need Domain-Specific Data Augmentations When Scaling Self-Supervised Learning
Théo Moutakanni, Maxime Oquab, Marc Szafraniec, Maria Vakalopoulou, Piotr Bojanowski
Abstract
Self-Supervised learning (SSL) with Joint-Embedding Architectures (JEA) has led to outstanding performances. All instantiations of this paradigm were trained using strong and well-established hand-crafted data augmentations, leading to the general belief that they are required for the proper training and performance of such models. On the other hand, generative reconstruction-based models such as BEIT and MAE or Joint-Embedding Predictive Architectures such as I-JEPA have shown strong performance without using data augmentations except masking. In this work, we challenge the importance of invariance and data-augmentation in JEAs at scale. By running a case-study on a recent SSL foundation model - DINOv2 - we show that strong image representations can be obtained with JEAs and only cropping without resizing provided the training data is large enough, reaching state-of-the-art results and using the least amount of augmentation in the literature. Through this study, we also discuss the impact of compute constraints on the outcomes of experimental deep learning research, showing that they can lead to very different conclusions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c27e3247-f4bb-4bcd-9cb9-94fc404068f7Cited by top-tier papers2
- Mosic: Optimal-Transport Motion Trajectory for Dense Self-Supervised LearningMohammadreza Salehi, Shashanka Venkataramanan, Ioana Simion, Efstratios Gavves et al.ICCV 2025
- A Cross Modal Knowledge Distillation & Data Augmentation Recipe for Improving Transcriptomics Representations through Morphological FeaturesIhab Bendidi, Yassir El Mesbahi, Alisandra Kaye Denton, Karush Suri et al.ICML 2025
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Self-Supervised Learning from Images with a Joint-Embedding Predictive ArchitectureMahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski et al.CVPR 2023
- Self-Supervised Learning Based on Transformed Image Reconstruction for Equivariance-Coherent Feature RepresentationQin Wang, Alessio Quercia, Benjamin Bruns, Abigail Morrison et al.AAAI 2026 · 2 citations
- T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular DataHugo Thimonier, José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel et al.ICLR 2025
- How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation NetworksEtai Littwin, Omid Saremi, Madhu Advani, Vimal Thilak et al.NeurIPS 2024 · 37 citations
- seq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World ModelsHafez Ghaemi, Eilif B. Muller, Shahab BakhtiariNeurIPS 2025 · 8 citations
