FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
Jiaqi Zhang, Juntuo Wang, Zhixin Sun, John Zou, Randall Balestriero
Abstract
Large-scale vision foundation models such as DINOv2 boast impressive performances by leveraging massive architectures and training datasets. But numerous scenarios require practitioners to reproduce those pre-training solutions, such as on private data, new modalities, or simply for scientific questioning-which is currently extremely demanding computation-wise. We thus propose a novel pretraining strategy for DINOv2 that simultaneously accelerates convergence-and strengthens robustness to common corruptions as a by-product. Our approach involves a frequency filtering curriculum-low-frequency being seen first-and the Gaussian noise patching augmentation. Applied to a ViT-B/16 backbone trained on ImageNet-1K, while pre-training time and FLOPs are reduced by 1.6× and 2.25×, our method still achieves matching robustness in corruption benchmarks (ImageNet-C) and maintains competitive linear probing performance compared with baseline. This dual benefit of efficiency and robustness makes large-scale self-supervised foundation modeling more attainable, while opening the door to novel exploration around data curriculum and augmentation as means to improve self-supervised learning models robustness. The code is available at https://github.com/KevinZ0217/fast_dinov2 Figure 1: The FastDINOv2 training pipeline comprises two stages. In the first stage, the initial 75% of training epochs utilize only low-frequency features extracted via downsampling. In the second stage, the remaining 25% of epochs employ full-resolution images with Gaussian noise patching.
Table 1: Comparison of training costs, evaluation on ImageNet-C, and linear probing on the ImageNet-1K validation set using a frozen ViT-B DINOv2 backbone trained on ImageNet-1K. Training Method Training Time (days)(↓) ImageNet-1K(↑) ImageNet-C (↓) GFLOPs (↓) DINOv2 16.64 (NVIDIA L40S) 77.8% 56.5% 493.76 FastDINOv2 10.32 (NVIDIA L40S) 76.2% 56.7% 219.92 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a85fb734-a950-4cc6-bd97-f11b063aec4eBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data CurriculumWenquan Lu, Jiaqi Zhang, Hugues Van Assel, Randall BalestrieroNeurIPS 2025 · 5 citations
- Accelerating Augmentation Invariance PretrainingJinhong Lin, Cheng-En Wu, Yibing Wei, Pedro MorgadoNeurIPS 2024 · 1 citation
- ExPLoRA: Parameter-Efficient Extended Pre-Training to Adapt Vision Transformers under Domain ShiftsSamar Khanna, Medhanie Irgau, David B. Lobell, Stefano ErmonICML 2025
- Beyond Random Augmentations: Pretraining with Hard ViewsFabio Ferreira, Ivo Rapant, Jörg K. H. Franke, Frank HutterICLR 2025
- Masked Frequency Modeling for Self-Supervised Visual Pre-TrainingJiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu et al.ICLR 2023 · 29 citations
