FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed
Jiaqi Zhang, Juntuo Wang, Zhixin Sun, John Zou, Randall Balestriero
摘要
Large-scale vision foundation models such as DINOv2 boast impressive performances by leveraging massive architectures and training datasets. But numerous scenarios require practitioners to reproduce those pre-training solutions, such as on private data, new modalities, or simply for scientific questioning-which is currently extremely demanding computation-wise. We thus propose a novel pretraining strategy for DINOv2 that simultaneously accelerates convergence-and strengthens robustness to common corruptions as a by-product. Our approach involves a frequency filtering curriculum-low-frequency being seen first-and the Gaussian noise patching augmentation. Applied to a ViT-B/16 backbone trained on ImageNet-1K, while pre-training time and FLOPs are reduced by 1.6× and 2.25×, our method still achieves matching robustness in corruption benchmarks (ImageNet-C) and maintains competitive linear probing performance compared with baseline. This dual benefit of efficiency and robustness makes large-scale self-supervised foundation modeling more attainable, while opening the door to novel exploration around data curriculum and augmentation as means to improve self-supervised learning models robustness. The code is available at https://github.com/KevinZ0217/fast_dinov2 Figure 1: The FastDINOv2 training pipeline comprises two stages. In the first stage, the initial 75% of training epochs utilize only low-frequency features extracted via downsampling. In the second stage, the remaining 25% of epochs employ full-resolution images with Gaussian noise patching.
Table 1: Comparison of training costs, evaluation on ImageNet-C, and linear probing on the ImageNet-1K validation set using a frozen ViT-B DINOv2 backbone trained on ImageNet-1K. Training Method Training Time (days)(↓) ImageNet-1K(↑) ImageNet-C (↓) GFLOPs (↓) DINOv2 16.64 (NVIDIA L40S) 77.8% 56.5% 493.76 FastDINOv2 10.32 (NVIDIA L40S) 76.2% 56.7% 219.92 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data CurriculumWenquan Lu, Jiaqi Zhang, Hugues Van Assel, Randall BalestrieroNeurIPS 2025 · 被引用 5 次
- Accelerating Augmentation Invariance PretrainingJinhong Lin, Cheng-En Wu, Yibing Wei, Pedro MorgadoNeurIPS 2024 · 被引用 1 次
- ExPLoRA: Parameter-Efficient Extended Pre-Training to Adapt Vision Transformers under Domain ShiftsSamar Khanna, Medhanie Irgau, David B. Lobell, Stefano ErmonICML 2025
- Beyond Random Augmentations: Pretraining with Hard ViewsFabio Ferreira, Ivo Rapant, Jörg K. H. Franke, Frank HutterICLR 2025
- Masked Frequency Modeling for Self-Supervised Visual Pre-TrainingJiahao Xie, Wei Li, Xiaohang Zhan, Ziwei Liu 等ICLR 2023 · 被引用 29 次
