Improving Robust Generalization by Direct PAC-Bayesian Bound Minimization
Zifan Wang, Nan Ding, Tomer Levinboim, Xi Chen, Radu Soricut
摘要
Recent research in robust optimization has shown an overfitting-like phenomenon in which models trained against adversarial attacks exhibit higher robustness on the training set compared to the test set. Although previous work provided theoretical explanations for this phenomenon using a robust PAC-Bayesian bound over the adversarial test error, related algorithmic derivations are at best only loosely connected to this bound, which implies that there is still a gap between their empirical success and our understanding of adversarial robustness theory. To close this gap, in this paper we consider a different form of the robust PAC-Bayesian bound and directly minimize it with respect to the model posterior. The derivation of the optimal solution connects PAC-Bayesian learning to the geometry of the robust loss surface through a Trace of Hessian (TrH) regularizer that measures the surface flatness. In practice, we restrict the TrH regularizer to the top layer only, which results in an analytical solution to the bound whose computational cost does not depend on the network depth. Finally, we evaluate our TrH regularization approach over CIFAR-10/100 and ImageNet using Vision Transformers (ViT) and compare against baseline adversarial robustness algorithms. Experimental results show that TrH regularization leads to improved ViT robustness that either matches or surpasses previous state-of-the-art approaches while at the same time requires less memory and computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Demystifying Structural Disparity in Graph Neural Networks: Can One Size Fit All?Haitao Mao, Zhikai Chen, Wei Jin, Haoyu Han 等NeurIPS 2023 · 被引用 58 次
- Variational Learning Finds Flatter Solutions at the Edge of StabilityAvrajit Ghosh, Bai Cong, Rio Yokota, Saiprasad Ravishankar 等NeurIPS 2025 · 被引用 2 次
- Efficient Source-Free Time-Series Adaptation via Parameter Subspace DisentanglementGaurav Patel, Christopher Michael Sandino, Behrooz Mahasseni, Ellen L. Zippi 等ICLR 2025
- CGU-Bayes: Causal Graph Uncertainty-Guided Bayesian Inference for Domain GeneralizationNaiyu Yin, Hanjing Wang, Yue Yu, Tian Gao 等CVPR 2026
它引用的顶会 Paper22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
相关 Paper
- CR-SAM: Curvature Regularized Sharpness-Aware MinimizationTao Wu, Tie Luo, Donald C. Wunsch IIAAAI 2024 · 被引用 15 次
- Benign Overfitting in Adversarial Training for Vision TransformersJiaming Zhang, Meng Ding, Shaopeng Fu, Jingfeng Zhang 等ICML 2026 · 被引用 1 次
- Relating Adversarially Robust Generalization to Flat MinimaDavid Stutz, Matthias Hein, Bernt SchieleICCV 2021 · 被引用 80 次
- Towards Robust Vision Transformers: Path Dependency Analysis and a Simple Two-Stage Adversarial TrainingSeongmin Kim, Byung Cheol SongCVPR 2026
- A PAC-Bayes Analysis of Adversarial RobustnessPaul Viallard, Guillaume Vidot, Amaury Habrard, Emilie MorvantNeurIPS 2021 · 被引用 21 次
