Adversarial Weight Perturbation Improves Generalization in Graph Neural Networks
Yihan Wu, Aleksandar Bojchevski, Heng Huang
Abstract
A lot of theoretical and empirical evidence shows that the flatter local minima tend to improve generalization. Adversarial Weight Perturbation (AWP) is an emerging technique to efficiently and effectively find such minima. In AWP we minimize the loss w.r.t. a bounded worst-case perturbation of the model parameters thereby favoring local minima with a small loss in a neighborhood around them. The benefits of AWP, and more generally the connections between flatness and generalization, have been extensively studied for i.i.d. data such as images. In this paper, we extensively study this phenomenon for graph data. Along the way, we first derive a generalization bound for non-i.i.d. node classification tasks. Then we identify a vanishing-gradient issue with all existing formulations of AWP and we propose a new Weighted Truncated AWP (WT-AWP) to alleviate this issue. We show that regularizing graph neural networks with WT-AWP consistently improves both natural and robust generalization across many different graph learning tasks and models. Our code is available at https://github.com/YihanWu95/WT-AWP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3879513-4ec1-4a14-b474-dc0305f1a855Cited by top-tier papers12
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 102 citations
- GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned ExpertsShirley Wu, Kaidi Cao, Bruno Ribeiro, James Y. Zou et al.NeurIPS 2024 · 27 citations
- PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement LearningJunfeng Guo, Ang Li, Lixu Wang, Cong LiuICCV 2023 · 27 citations
- Defending against Data-Free Model Extraction by Distributionally Robust Defensive TrainingZhenyi Wang, Li Shen, Tongliang Liu, Tiehang Duan et al.NeurIPS 2023 · 26 citations
- Solving a Class of Non-Convex Minimax Optimization in Federated LearningXidong Wu, Jianhui Sun, Zhengmian Hu, Aidong Zhang et al.NeurIPS 2023 · 26 citations
Builds on12
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
- A PAC-Bayesian Approach to Generalization Bounds for Graph Neural NetworksRenjie Liao, Raquel Urtasun, Richard S. ZemelICLR 2021 · 109 citations
- Subgroup Generalization and Fairness of Graph Neural NetworksJiaqi Ma, Junwei Deng, Qiaozhu MeiNeurIPS 2021 · 102 citations
Related papers
- Regularizing Neural Networks via Adversarial Model PerturbationYaowei Zheng, Richong Zhang, Yongyi MaoCVPR 2021
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Robust Optimization as Data Augmentation for Large-scale GraphsKezhi Kong, Guohao Li, Mucong Ding, Zuxuan Wu et al.CVPR 2022 · 87 citations
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou et al.CVPR 2023
- Lipschitz Bounds and Provably Robust Training by Laplacian SmoothingVishaal Krishnan, Abed AlRahman Al Makdah, Fabio PasqualettiNeurIPS 2020 · 28 citations
