ICML2026

AdaS: Adaptive Gradient Descent for Spiking Transformers

Zijian Zhou, Honglin Cao, Ammar Belatreche, Wenjie Wei, Yimeng Shan, Yu Liang, Yu Yang, Shuai Wang, Yalan Ye, Malu Zhang, Yang Yang, Haizhou Li

Abstract

Transformer-based Spiking Neural Networks (SNNs) combine Transformer performance with SNN energy efficiency through an event-driven self-attention mechanism. However, Spiking Transformers still lag behind their Artificial Neural Network (ANN) counterparts. Most existing studies address this issue through new architectural designs, yet few have explored optimization algorithms tailored to Spiking Transformers. We substantiate the excessive noise problem in Spiking Transformer training by quantitatively defining parameter-update noise and, based on this definition, providing theoretical analysis and experimental validation. To address this problem, we propose AdaS, an adaptive gradient descent method for Spiking Transformers. AdaS reduces excessive noise by adaptively incorporating a gradient update component into adaptive optimization. Instead of simply removing noise, AdaS maintains it at an appropriate level to preserve its generalization benefits, thereby improving the performance of Spiking Transformers. We conduct extensive experiments on various Spiking Transformer architectures and datasets from both computer vision and natural language processing. The results demonstrate that the proposed AdaS consistently enhances performance across different Spiking Transformers, validating its effectiveness and generalizability. This work is among the first systematic studies of optimizer design specifically for Spiking Transformers, offering a practical tool to narrow the accuracy gap with ANNs while preserving the energy advantages of spike-based computation. Code is available at https://github.com/CayleyZ/AdaS.