Spiking Transformers for Event-based Single Object Tracking
Jiqing Zhang, Bo Dong, Haiwei Zhang, Jianchuan Ding, Felix Heide, Baocai Yin, Xin Yang
Abstract
Event-based cameras bring a unique capability to tracking, being able to function in challenging real-world conditions as a direct result of their high temporal resolution and high dynamic range. These imagers capture events asynchronously that encode rich temporal and spatial information. However, effectively extracting this information from events remains an open challenge. In this work, we propose a spiking transformer network, STNet, for single object tracking. STNet dynamically extracts and fuses information from both temporal and spatial domains. In particular, the proposed architecture features a transformer module to provide global spatial information and a spiking neural network (SNN) module for extracting temporal cues. The spiking threshold of the SNN module is dynamically adjusted based on the statistical cues of the spatial information, which we find essential in providing robust SNN features. We fuse both feature branches dynamically with a novel cross-domain attention fusion algorithm. Extensive experiments on three event-based datasets, FE240hz, EED and VisEvent validate that the proposed STNet outperforms existing state-of-the-art methods in both tracking accuracy and speed with a significant margin. The code and pretrained models are at https://github.com/Jee- King/CVPR2022_STNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51274ebe-9b05-4d8b-aff0-e38eb25b2ec2Cited by top-tier papers55
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
- Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic ChipsMan Yao, Jiakui Hu, Tianxiang Hu, Yifan Xu et al.ICLR 2024 · 154 citations
- Deep Directly-Trained Spiking Neural Networks for Object DetectionQiaoyi Su, Yuhong Chou, Yifan Hu, Jianing Li et al.ICCV 2023 · 143 citations
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun et al.ICCV 2023 · 86 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Learning Discriminative Model Prediction for TrackingGoutam Bhat, Martin Danelljan, Luc Van Gool, Radu TimofteICCV 2019 · 1,294 citations
- Learning Spatio-Temporal Transformer for Visual TrackingBin Yan, Houwen Peng, Jianlong Fu, Dong Wang et al.ICCV 2021 · 1,062 citations
- SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation GuidelinesYinda Xu, Zeyu Wang, Zuoxin Li, Ye Yuan et al.AAAI 2020 · 944 citations
- GradNet: Gradient-Guided Network for Visual Object TrackingPeixia Li, Boyu Chen, Wanli Ouyang, Dong Wang et al.ICCV 2019 · 255 citations
Related papers
- SDTrack: A Baseline for Event-based Tracking via Spiking Neural NetworksYimeng Shan, Zhenbang Ren, Haodi Wu, Wenjie Wei et al.CVPR 2026 · 14 citations
- SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural NetworkYang Wang, Jiqing Zhang, Chuanyu Sun, Qianhui Liu et al.CVPR 2026
- Fully Spiking Neural Networks for Unified Frame-Event Object TrackingJingjun Yang, Liangwei Fan, Jinpu Zhang, Xiangkai Lian et al.NeurIPS 2025 · 9 citations
- Hybrid Spiking Vision Transformer for Object Detection with Event CamerasQi Xu, Jie Deng, Jiangrong Shen, Biwu Chen et al.ICML 2025
- Event Stream Super-Resolution via Spatiotemporal Constraint LearningSiqi Li, Yutong Feng, Yipeng Li, Yu Jiang et al.ICCV 2021 · 25 citations
