Finding Visual Saliency in Continuous Spike Stream
Lin Zhu, Xianzhang Chen, Xiao Wang, Hua Huang
Abstract
As a bio-inspired vision sensor, the spike camera emulates the operational principles of the fovea, a compact retinal region, by employing spike discharges to encode the accumulation of per-pixel luminance intensity. Leveraging its high temporal resolution and bio-inspired neuromorphic design, the spike camera holds significant promise for advancing computer vision applications. Saliency detection mimics the behavior of human beings and captures the most salient region from the scenes. In this paper, we investigate the visual saliency in the continuous spike stream for the first time. To effectively process the binary spike stream, we propose a Recurrent Spiking Transformer (RST) framework, which is based on a full spiking neural network. Our framework enables the extraction of spatio-temporal features from the continuous spatio-temporal spike stream while maintaining low power consumption. To facilitate the training and validation of our proposed model, we build a comprehensive real-world spike-based visual saliency dataset, enriched with numerous light conditions. Extensive experiments demonstrate the superior performance of our Recurrent Spiking Transformer framework in comparison to other spike neural network-based methods. Our framework exhibits a substantial margin of improvement in capturing and highlighting visual saliency in the spike stream, which not only provides a new perspective for spike-based saliency segmentation but also shows a new paradigm for full SNN-based transformer models. The code and dataset are available at https://github.com/BIT-Vision/SVS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7725d748-4e8c-4fd1-a5aa-ede47a51159dCited by top-tier papers1
Ask how each one uses itBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Spiking-YOLO: Spiking Neural Network for Energy-Efficient Object DetectionSei Joon Kim, Seongsik Park, Byunggook Na, Sungroh YoonAAAI 2020 · 512 citations
- Global Context-Aware Progressive Aggregation Network for Salient Object DetectionZuyao Chen, Qianqian Xu, Runmin Cong, Qingming HuangAAAI 2020 · 481 citations
- Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural NetworksJesse J. Hagenaars, Federico Paredes-Vallés, Guido de CroonNeurIPS 2021 · 178 citations
Related papers
- Retina-Like Visual Image Reconstruction via Spiking Neural ModelLin Zhu, Siwei Dong, Jianing Li, Tiejun Huang et al.CVPR 2020
- Recurrent Spike-based Image Restoration under General IlluminationLin Zhu, Yunlong Zheng, Mengyue Geng, Lizhi Wang et al.ACM MM 2023 · 9 citations
- Spk2VidNet: A Hierarchical Recurrent Architecture for High-Fidelity Video Reconstruction from Long Spike-Camera StreamsYuanlin Wang, Ruiqin Xiong, Jiyu Xie, Zhenkun Zhu et al.CVPR 2026
- Learning Optical Flow from Continuous Spike StreamsRui Zhao, Ruiqin Xiong, Jing Zhao, Zhaofei Yu et al.NeurIPS 2022 · 49 citations
- Learning to Super-resolve Dynamic Scenes for Neuromorphic Spike CameraJing Zhao, Ruiqin Xiong, Jian Zhang, Rui Zhao et al.AAAI 2023 · 21 citations
