Efficient Sharpness-Aware Minimization for Molecular Graph Transformer Models
Yili Wang, Kaixiong Zhou, Ninghao Liu, Ying Wang, Xin Wang
摘要
Sharpness-aware minimization (SAM) has received increasing attention in computer vision since it can effectively eliminate the sharp local minima from the training trajectory and mitigate generalization degradation. However, SAM requires two sequential gradient computations during the optimization of each step: one to obtain the perturbation gradient and the other to obtain the updating gradient. Compared with the base optimizer (e.g., Adam), SAM doubles the time overhead due to the additional perturbation gradient. By dissecting the theory of SAM and observing the training gradient of the molecular graph transformer, we propose a new algorithm named GraphSAM, which reduces the training cost of SAM and improves the generalization performance of graph transformer models. There are two key factors that contribute to this result: (i) gradient approximation: we use the updating gradient of the previous step to approximate the perturbation gradient at the intermediate steps smoothly (increases efficiency); (ii) loss landscape approximation: we theoretically prove that the loss landscape of GraphSAM is limited to a small range centered on the expected loss of SAM (guarantees generalization performance). The extensive experiments on six datasets with different tasks demonstrate the superiority of GraphSAM, especially in optimizing the model update process. The code is in: https://github.com/YL-wang/GraphSAM/tree/graphsam . Recently, sharpness aware minimization (SAM) (Foret et al., 2020) has been proposed to explicitly smooth the sharp local minima during model training like pre-training. Nevertheless, SAM requires two forward and backward propagations at each step: one to obtain the worst-case adversarial gradient
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Dynamic Graph Unlearning: A General and Efficient Post-Processing Method via Gradient TransformationHe Zhang, Bang Wu, Xiangwen Yang, Xingliang Yuan 等WWW 2025 · 被引用 16 次
- Optimizing OOD Detection in Molecular Graphs: A Novel Approach with Diffusion ModelsXu Shen, Yili Wang, Kaixiong Zhou, Shirui Pan 等KDD 2024 · 被引用 12 次
- Rethinking Independent Cross-Entropy Loss For Graph-Structured DataRui Miao, Kaixiong Zhou, Yili Wang, Ninghao Liu 等ICML 2024 · 被引用 5 次
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 被引用 2 次
- Unifying Unsupervised Graph-Level Anomaly Detection and Out-of-Distribution Detection: A BenchmarkYili Wang, Yixin Liu, Xu Shen, Chenyu Li 等ICLR 2025 · 被引用 2 次
它引用的顶会 Paper25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
相关 Paper
- Fast Graph Sharpness-Aware Minimization for Enhancing and Accelerating Few-Shot Node ClassificationYihong Luo, Yuhan Chen, Siya Qiu, Yiwei Wang 等NeurIPS 2024 · 被引用 7 次
- Towards Efficient and Scalable Sharpness-Aware MinimizationYong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh 等CVPR 2022 · 被引用 61 次
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou 等ICLR 2022 · 被引用 168 次
- Random Sharpness-Aware MinimizationYong Liu, Siqi Mai, Minhao Cheng, Xiangning Chen 等NeurIPS 2022 · 被引用 38 次
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 被引用 16 次
