Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
Brian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, Bhavya Kailkhura
摘要
This paper revisits the simple, long-studied, yet still unsolved problem of making image classifiers robust to imperceptible perturbations. Taking CIFAR10 as an example, SOTA clean accuracy is about %, but SOTA robustness to -norm bounded perturbations barely exceeds %. To understand this gap, we analyze how model size, dataset size, and synthetic data quality affect robustness by developing the first scaling laws for adversarial training. Our scaling laws reveal inefficiencies in prior art and provide actionable feedback to advance the field. For instance, we discovered that SOTA methods diverge notably from compute-optimal setups, using excess compute for their level of robustness. Leveraging a compute-efficient setup, we surpass the prior SOTA with % (%) fewer training (inference) FLOPs. We trained various compute-efficient models, with our best achieving % AutoAttack accuracy (% gain). However, our scaling laws also predict robustness slowly grows then plateaus at %: dwarfing our new SOTA by scaling is impractical, and perfect robustness is impossible. To better understand this predicted limit, we carry out a small-scale human evaluation on the AutoAttack data that fools our top-performing model. Concerningly, we estimate that human performance also plateaus near %, which we show to be attributable to -constrained attacks' generation of invalid images not consistent with their original labels. Having characterized limiting roadblocks, we outline promising paths for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Robustness in Both Domains: CLIP Needs a Robust Text EncoderElías Abad-Rocamora, Christian Schlarmann, Naman Deep Singh, Yongtao Wu 等NeurIPS 2025 · 被引用 4 次
- Adversarial Vulnerability from Interference Between Features in SuperpositionEdward Stevinson, Lucas Prieto, Melih Barsbey, Tolga BirdalICML 2026 · 被引用 4 次
- Monitoring Robustness and Individual FairnessAshutosh Gupta, Thomas A. Henzinger, Konstantin Kueffner, Kaushik Mallik 等KDD 2025 · 被引用 2 次
- Scaling and Taming Adversarial Training with Synthetic DataJuntao Wu, Xianting Huang, Yu Chen, Shuai Pang 等ICCV 2025 · 被引用 2 次
- Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial RobustnessJuntao Wu, Ziyu Song, Xiaoyu Zhang, Shujun Xie 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
相关 Paper
- Revisiting Adversarial Training at ScaleZeyu Wang, Xianhang Li, Hongru Zhu, Cihang XieCVPR 2024
- Revisiting Residual Networks for Adversarial RobustnessShihua Huang, Zhichao Lu, Kalyanmoy Deb, Vishnu Naresh BoddetiCVPR 2023
- Scaling Trends in Language Model RobustnessNikolaus H. R. Howe, Ian R. McKenzie, Oskar John Hollinsworth, Michal Zajac 等ICML 2025
- Data Augmentation Can Improve RobustnessSylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg 等NeurIPS 2021 · 被引用 427 次
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg 等NeurIPS 2021 · 被引用 384 次
