Lune

NeurIPS2025Top-tier venue

FastDINOv2: Frequency Based Curriculum Learning Improves Robustness and Training Speed

Jiaqi Zhang, Juntuo Wang, Zhixin Sun, John Zou, Randall Balestriero

2025Year
5Citations

Abstract

Large-scale vision foundation models such as DINOv2 boast impressive performances by leveraging massive architectures and training datasets. But numerous scenarios require practitioners to reproduce those pre-training solutions, such as on private data, new modalities, or simply for scientific questioning-which is currently extremely demanding computation-wise. We thus propose a novel pretraining strategy for DINOv2 that simultaneously accelerates convergence-and strengthens robustness to common corruptions as a by-product. Our approach involves a frequency filtering curriculum-low-frequency being seen first-and the Gaussian noise patching augmentation. Applied to a ViT-B/16 backbone trained on ImageNet-1K, while pre-training time and FLOPs are reduced by 1.6× and 2.25×, our method still achieves matching robustness in corruption benchmarks (ImageNet-C) and maintains competitive linear probing performance compared with baseline. This dual benefit of efficiency and robustness makes large-scale self-supervised foundation modeling more attainable, while opening the door to novel exploration around data curriculum and augmentation as means to improve self-supervised learning models robustness. The code is available at https://github.com/KevinZ0217/fast_dinov2 Figure 1: The FastDINOv2 training pipeline comprises two stages. In the first stage, the initial 75% of training epochs utilize only low-frequency features extracted via downsampling. In the second stage, the remaining 25% of epochs employ full-resolution images with Gaussian noise patching.

Table 1: Comparison of training costs, evaluation on ImageNet-C, and linear probing on the ImageNet-1K validation set using a frozen ViT-B DINOv2 backbone trained on ImageNet-1K. Training Method Training Time (days)(↓) ImageNet-1K(↑) ImageNet-C (↓) GFLOPs (↓) DINOv2 16.64 (NVIDIA L40S) 77.8% 56.5% 493.76 FastDINOv2 10.32 (NVIDIA L40S) 76.2% 56.7% 219.92 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a85fb734-a950-4cc6-bd97-f11b063aec4e

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines