ALERT-Transformer: Bridging Asynchronous and Synchronous Machine Learning for Real-Time Event-based Spatio-Temporal Data
Carmen Martin-Turrero, Maxence Bouvier, Manuel Breitenstein, Pietro Zanuttigh, Vincent Parret
Abstract
We seek to enable classic processing of continuous ultra-sparse spatiotemporal data generated by event-based sensors with dense machine learning models. We propose a novel hybrid pipeline composed of asynchronous sensing and synchronous processing that combines several ideas: (1) an embedding based on PointNet models -- the ALERT module -- that can continuously integrate new and dismiss old events thanks to a leakage mechanism, (2) a flexible readout of the embedded data that allows to feed any downstream model with always up-to-date features at any sampling rate, (3) exploiting the input sparsity in a patch-based approach inspired by Vision Transformer to optimize the efficiency of the method. These embeddings are then processed by a transformer model trained for object and gesture recognition. Using this approach, we achieve performances at the state-of-the-art with a lower latency than competitors. We also demonstrate that our asynchronous model can operate at any desired sampling rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Maximizing Asynchronicity in Event-based Neural NetworksHaiqing Hao, Nikola Zubic, Weihua He, Zhipeng Sui et al.ICLR 2026 · 2 citations
- Synthetic Series-Symbol Data Generation for Time Series Foundation ModelsWenxuan Wang, Kai Wu, Yujian Betterest Li, Dan Wang et al.NeurIPS 2025 · 1 citation
- Learning to Match Unpaired Data with Minimum Entropy CouplingMustapha Bounoua, Giulio Franzese, Pietro MichiardiICML 2025
- FLAME: Fast Long-context Adaptive Memory for Event-based VisionBiswadeep Chakraborty, Saibal MukhopadhyayNeurIPS 2025
- S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision ProcessingPing He, Rong Xiao, Wanying Xu, Chenwei Tang et al.AAAI 2026
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- End-to-End Learning of Representations for Asynchronous Event-Based DataDaniel Gehrig, Antonio Loquercio, Konstantinos G. Derpanis, Davide ScaramuzzaICCV 2019 · 427 citations
- Point TransformerHengshuang Zhao, Li Jiang, Jiaya Jia, Philip H. S. Torr et al.ICCV 2021 · 23 citations
Related papers
- AEGNN: Asynchronous Event-based Graph Neural NetworksSimon Schaefer, Daniel Gehrig, Davide ScaramuzzaCVPR 2022 · 135 citations
- Asynchronous Event Processing with Local-Shift Graph Convolutional NetworkLinhui Sun, Yifan Zhang, Jian Cheng, Hanqing LuAAAI 2023 · 2 citations
- Rethinking Scale-Aware Temporal Encoding for Event-based Object DetectionLin Zhu, Tengyu Long, Xiao Wang, Lizhi Wang et al.NeurIPS 2025 · 4 citations
- Event-based Video Reconstruction Using TransformerWenming Weng, Yueyi Zhang, Zhiwei XiongICCV 2021 · 139 citations
- TTPOINT: A Tensorized Point Cloud Network for Lightweight Action Recognition with Event CamerasHongwei Ren, Yue Zhou, Haotian Fu, Yulong Huang et al.ACM MM 2023 · 14 citations
