Otters: An Energy-Efficient Spiking Transformer via Optical Time-to-First-Spike Encoding
Zhanglu Yan, Jiayi Mao, Qianhui Liu, Fanfan Li, Tao Luo, Gang Pan, Bowen Zhu, Weng-Fai Wong
Abstract
Spiking neural networks (SNNs) promise high energy efficiency, particularly with time-to-first-spike (TTFS) encoding, which maximizes sparsity by emitting at most one spike per neuron. However, this energy advantage is often unrealized because inference requires evaluating a temporal decay function and then multiplying the result by the synaptic weights. This paper challenges this costly approach by repurposing a physical hardware `bug', namely, the natural signal decay in optoelectronic devices, as the core computation of TTFS. We fabricated a custom indium oxide optoelectronic synapse that demonstrates how its intrinsic physical decay directly implements the required temporal function. By treating the device's analog output as the fused product of the synaptic weight and temporal decay, optoelectronic synaptic TTFS (named Otters) eliminates these expensive digital operations. To use the Otters' paradigm in complex architectures such as the transformer, which are challenging to train directly due to sparsity, we introduce a novel quantized neural network-to-SNN conversion algorithm. This complete hardware-software co-design enables our model to achieve state-of-the-art accuracy across seven GLUE benchmark datasets and demonstrates a 1.77 improvement in energy efficiency over previous leading SNNs, based on a comprehensive analysis of compute, data movement, and memory access costs using energy measurements from a commercial 22nm process. Our work thus establishes a new paradigm for energy-efficient SNNs that translates fundamental device physics directly into powerful computational primitives. All codes and data are open source.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- BiT: Robustly Binarized Multi-distilled TransformerZechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao et al.NeurIPS 2022 · 93 citations
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit DifferentiationMalyaban Bal, Abhronil SenguptaAAAI 2024 · 78 citations
- Temporal-Coded Spiking Neural Networks with Dynamic Firing Threshold: Learning with Event-Driven BackpropagationWenjie Wei, Malu Zhang, Hong Qu, Ammar Belatreche et al.ICCV 2023 · 41 citations
- Efficiently Training Time-to-First-Spike Spiking Neural Networks from ScratchKaiwei Che, Wei Fang, Zhengyu Ma, Yifan Huang et al.ICML 2026 · 3 citations
Related papers
- Training-Free ANN-to-SNN Conversion for High-Performance Spiking TransformersJingya Wang, Xin Deng, Wenjie Wei, Dehao Zhang et al.AAAI 2026 · 1 citation
- TTFSFormer: A TTFS-based Lossless Conversion of Spiking TransformerLusen Zhao, Zihan Huang, Jianhao Ding, Zhaofei YuICML 2025
- A time-to-first-spike coding and conversion aware training for energy-efficient deep spiking neural network processor designDongwoo Lew, Kyungchul Lee, Jongsun ParkDAC 2022 · 18 citations
- Parallel Training Time-to-First-Spike Spiking Neural NetworksKaiwei Che, Wei Fang, Peng Xue, Yifan Huang et al.AAAI 2026
- Temporal-coded Spiking TransformerQian Sun, Chengzhuo Lu, Wenyu Chen, Wenjie Wei et al.ACM MM 2025 · 1 citation
