EGSST: Event-based Graph Spatiotemporal Sensitive Transformer for Object Detection
Sheng Wu, Hang Sheng, Hui Feng, Bo Hu
摘要
Event cameras provide exceptionally high temporal resolution in dynamic vision systems due to their unique event-driven mechanism. However, the sparse and asynchronous nature of event data makes frame-based visual processing methods inappropriate. This study proposes a novel framework, Event-based Graph Spa-tiotemporal Sensitive Transformer (EGSST), for the exploitation of spatial and temporal properties of event data. Firstly, a well-designed graph structure is employed to model event data, which not only preserves the original temporal data but also captures spatial details. Furthermore, inspired by the phenomenon that human eyes pay more attention to objects that produce significant dynamic changes, we design a Spatiotemporal Sensitivity Module (SSM) and an adaptive Temporal Activation Controller (TAC). Through these two modules, our framework can mimic the response of the human eyes in dynamic environments by selectively activating the temporal attention mechanism based on the relative dynamics of event data, thereby effectively conserving computational resources. In addition, the integration of a lightweight, multi-scale Linear Vision Transformer (LViT) markedly enhances processing efficiency. Our research proposes a fully event-driven approach, effectively exploiting the temporal precision of event data and optimising the allocation of computational resources by intelligently distinguishing the dynamics within the event data. The framework provides a lightweight, fast, accurate, and fully event-based solution for object detection tasks in complex dynamic environments, demonstrating significant practicality and potential for application. The source code can be found at: EGSST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EventFlash: Towards Efficient MLLMs for Event-Based VisionShaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang 等ICLR 2026 · 被引用 5 次
- SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action RecognitionRui Fan, Weidong Hao, Juntao Guan, Lai Rui 等CVPR 2026 · 被引用 1 次
- Seeing Motion Through Polarity for Event-based Action RecognitionMeiqi Cao, Jiachao Zhang, Xin Jiang, Rui Yan 等CVPR 2026
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
相关 Paper
- Scene Adaptive Sparse Transformer for Event-based Object DetectionYansong Peng, Hebei Li, Yueyi Zhang, Xiaoyan Sun 等CVPR 2024 · 被引用 25 次
- Eye-based Emotion Recognition via Event-Driven Sparse TransformersZixuan Wan, Jiqing Zhang, Yushan Wang, Hu Lin 等ACM MM 2025 · 被引用 2 次
- Adaptive Vision Transformer for Event-Based Human Pose EstimationNannan Yu, Tao Ma, Jiqing Zhang, Yuji Zhang 等ACM MM 2024 · 被引用 9 次
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang 等NeurIPS 2025 · 被引用 4 次
- GET: Group Event Transformer for Event-Based VisionYansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun 等ICCV 2023 · 被引用 86 次
