Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
Brian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, Bhavya Kailkhura
Abstract
This paper revisits the simple, long-studied, yet still unsolved problem of making image classifiers robust to imperceptible perturbations. Taking CIFAR10 as an example, SOTA clean accuracy is about %, but SOTA robustness to -norm bounded perturbations barely exceeds %. To understand this gap, we analyze how model size, dataset size, and synthetic data quality affect robustness by developing the first scaling laws for adversarial training. Our scaling laws reveal inefficiencies in prior art and provide actionable feedback to advance the field. For instance, we discovered that SOTA methods diverge notably from compute-optimal setups, using excess compute for their level of robustness. Leveraging a compute-efficient setup, we surpass the prior SOTA with % (%) fewer training (inference) FLOPs. We trained various compute-efficient models, with our best achieving % AutoAttack accuracy (% gain). However, our scaling laws also predict robustness slowly grows then plateaus at %: dwarfing our new SOTA by scaling is impractical, and perfect robustness is impossible. To better understand this predicted limit, we carry out a small-scale human evaluation on the AutoAttack data that fools our top-performing model. Concerningly, we estimate that human performance also plateaus near %, which we show to be attributable to -constrained attacks' generation of invalid images not consistent with their original labels. Having characterized limiting roadblocks, we outline promising paths for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 942fa3b9-88be-4ff1-bd93-b1289eb517beCited by top-tier papers16
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu et al.NeurIPS 2025 · 4 citations
- Adversarial Vulnerability from Interference Between Features in SuperpositionEdward Stevinson, Lucas Prieto, Melih Barsbey, Tolga BirdalICML 2026 · 4 citations
- Monitoring Robustness and Individual FairnessAshutosh Gupta, Thomas A. Henzinger, Konstantin Kueffner, Kaushik Mallik et al.KDD 2025 · 2 citations
- Scaling and Taming Adversarial Training with Synthetic DataJuntao Wu, Xianting Huang, Yu Chen, Shuai Pang et al.ICCV 2025 · 2 citations
- Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial RobustnessJuntao Wu, Ziyu Song, Xiaoyu Zhang, Shujun Xie et al.AAAI 2025 · 1 citation
Builds on18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
Related papers
- Revisiting Adversarial Training at ScaleZeyu Wang, Xianhang Li, Hongru Zhu, Cihang XieCVPR 2024
- Revisiting Residual Networks for Adversarial RobustnessShihua Huang, Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh BoddetiCVPR 2023
- Scaling Trends in Language Model RobustnessNikolaus H. R. Howe, Ian R. McKenzie, Oskar John Hollinsworth, Michal Zajac et al.ICML 2025
- Data Augmentation Can Improve RobustnessSylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg et al.NeurIPS 2021 · 427 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
