Efficient Sharpness-Aware Minimization for Molecular Graph Transformer Models
Yili Wang, Kaixiong Zhou, Ninghao Liu, Ying Wang, Xin Wang
Abstract
Sharpness-aware minimization (SAM) has received increasing attention in computer vision since it can effectively eliminate the sharp local minima from the training trajectory and mitigate generalization degradation. However, SAM requires two sequential gradient computations during the optimization of each step: one to obtain the perturbation gradient and the other to obtain the updating gradient. Compared with the base optimizer (e.g., Adam), SAM doubles the time overhead due to the additional perturbation gradient. By dissecting the theory of SAM and observing the training gradient of the molecular graph transformer, we propose a new algorithm named GraphSAM, which reduces the training cost of SAM and improves the generalization performance of graph transformer models. There are two key factors that contribute to this result: (i) gradient approximation: we use the updating gradient of the previous step to approximate the perturbation gradient at the intermediate steps smoothly (increases efficiency); (ii) loss landscape approximation: we theoretically prove that the loss landscape of GraphSAM is limited to a small range centered on the expected loss of SAM (guarantees generalization performance). The extensive experiments on six datasets with different tasks demonstrate the superiority of GraphSAM, especially in optimizing the model update process. The code is in: https://github.com/YL-wang/GraphSAM/tree/graphsam . Recently, sharpness aware minimization (SAM) (Foret et al., 2020) has been proposed to explicitly smooth the sharp local minima during model training like pre-training. Nevertheless, SAM requires two forward and backward propagations at each step: one to obtain the worst-case adversarial gradient
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a596c21-82d5-4fd9-a6e2-6d24d9f176d6Cited by top-tier papers8
- Dynamic Graph Unlearning: A General and Efficient Post-Processing Method via Gradient TransformationHe Zhang, Bang Wu, Xiangwen Yang, Xingliang Yuan et al.WWW 2025 · 16 citations
- Optimizing OOD Detection in Molecular Graphs: A Novel Approach with Diffusion ModelsXu Shen, Yili Wang, Kaixiong Zhou, Shirui Pan et al.KDD 2024 · 12 citations
- Rethinking Independent Cross-Entropy Loss For Graph-Structured DataRui Miao, Kaixiong Zhou, Yili Wang, Ninghao Liu et al.ICML 2024 · 5 citations
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 2 citations
- Unifying Unsupervised Graph-Level Anomaly Detection and Out-of-Distribution Detection: A BenchmarkYili Wang, Yixin Liu, Xu Shen, Chenyu Li et al.ICLR 2025 · 2 citations
Builds on25
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- Fast Graph Sharpness-Aware Minimization for Enhancing and Accelerating Few-Shot Node ClassificationYihong Luo, Yuhan Chen, Siya Qiu, Yiwei Wang et al.NeurIPS 2024 · 7 citations
- Towards Efficient and Scalable Sharpness-Aware MinimizationYong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh et al.CVPR 2022 · 61 citations
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou et al.ICLR 2022 · 168 citations
- Random Sharpness-Aware MinimizationYong Liu, Siqi Mai, Minhao Cheng, Xiangning Chen et al.NeurIPS 2022 · 38 citations
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 16 citations
