Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models
Andy Zhou, Jindong Wang, Yu-Xiong Wang, Haohan Wang
Abstract
We propose a conceptually simple and lightweight framework for improving the robustness of vision models through the combination of knowledge distillation and data augmentation. We address the conjecture that larger models do not make for better teachers by showing strong gains in out-of-distribution robustness when distilling from pretrained foundation models. Following this finding, we propose Discrete Adversarial Distillation (DAD), which leverages a robust teacher to generate adversarial examples and a VQGAN to discretize them, creating more informative samples than standard data augmentation techniques. We provide a theoretical framework for the use of a robust teacher in the knowledge distillation with data augmentation setting and demonstrate strong gains in out-of-distribution robustness and clean accuracy across different student architectures. Notably, our method adds minor computational overhead compared to similar techniques and can be easily combined with other data augmentations for further improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c222b553-79f4-4417-b611-34cd39d8cd5aCited by top-tier papers4
- Can One Modality Model Synergize Training of Other Modality Models?Jae-Jun Lee, Sung Whan YoonICLR 2025
- OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation TriadLuyao Tang, Yuxuan Yuan, Chaoqi Chen, Zeyu Zhang et al.CVPR 2025
- TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization AbilityFengji Ma, Hei Victor Cheng, Chenxing Li, Li LiuAAAI 2026
- Open-Vocabulary Customization from CLIP via Data-Free Knowledge DistillationYongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang et al.ICLR 2025
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- Confound from all Sides, Distill with Resilience: Multi-Objective Adversarial Paths to Zero-Shot RobustnessJunhao Dong, Jiao Liu, Xinghua Qu, Yew-Soon OngICCV 2025 · 1 citation
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 258 citations
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student BetterBojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang JiangICCV 2021 · 136 citations
- Teacher as a Lenient Expert: Teacher-Agnostic Data-Free Knowledge DistillationHyunjune Shin, Dong-Wan ChoiAAAI 2024 · 8 citations
- Boosting Accuracy and Robustness of Student Models via Adaptive Adversarial DistillationBo Huang, Mingyang Chen, Yi Wang, Junda Lu et al.CVPR 2023
