GET: Group Event Transformer for Event-Based Vision
Yansong Peng, Yueyi Zhang, Zhiwei Xiong, Xiaoyan Sun, Feng Wu
摘要
Event cameras are a type of novel neuromorphic sensor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity. To address this issue, we propose a novel Group-based vision Transformer backbone for Event-based vision, called Group Event Transformer (GET), which decouples temporal-polarity information from spatial information throughout the feature extraction process. Specifically, we first propose a new event representation for GET, named Group Token, which groups asynchronous events based on their timestamps and polarities. Then, GET applies the Event Dual Self-Attention block, and Group Token Aggregation module to facilitate effective feature communication and integration in both the spatial and temporalpolarity domains. After that, GET can be integrated with different downstream tasks by connecting it with various heads. We evaluate our method on four event-based classification datasets (Cifar10-DVS, N-MNIST, N-CARS, and DVS128Gesture) and two event-based object detection datasets (1Mpx and Gen1), and the results demonstrate that GET outperforms other state-of-the-art methods. The code is available at https://github.com/Peterande/ GET-Group-Event-Transformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Single-Model and Any-Modality for Video Object TrackingZongwei Wu, Jilai Zheng, Xiangxuan Ren, Florin-Alexandru Vasluianu 等CVPR 2024 · 被引用 78 次
- State Space Models for Event CamerasNikola Zubic, Mathias Gehrig, Davide ScaramuzzaCVPR 2024 · 被引用 33 次
- Scene Adaptive Sparse Transformer for Event-based Object DetectionYansong Peng, Hebei Li, Yueyi Zhang, Xiaoyan Sun 等CVPR 2024 · 被引用 25 次
- SMamba: Sparse Mamba for Event-based Object DetectionNan Yang, Yang Wang, Zhanwen Liu, Meng Li 等AAAI 2025 · 被引用 17 次
- Event-Based Tiny Object Detection: A Benchmark Dataset and BaselineNuo Chen, Chao Xiao, Yimian Dai, Shiman He 等ICCV 2025 · 被引用 10 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
相关 Paper
- Recurrent Vision Transformers for Object Detection with Event CamerasMathias Gehrig, Davide ScaramuzzaCVPR 2023
- Dual Memory Aggregation Network for Event-Based Object Detection with Learnable RepresentationDongsheng Wang, Xu Jia, Yang Zhang, Xinyu Zhang 等AAAI 2023 · 被引用 22 次
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 被引用 139 次
- Better and Faster: Adaptive Event Conversion for Event-Based Object DetectionYansong Peng, Yueyi Zhang, Peilin Xiao, Xiaoyan Sun 等AAAI 2023 · 被引用 25 次
- Adaptive Vision Transformer for Event-Based Human Pose EstimationNannan Yu, Tao Ma, Jiqing Zhang, Yuji Zhang 等ACM MM 2024 · 被引用 9 次
