TESPEC: Temporally-Enhanced Self-Supervised Pretraining for Event Cameras
Mohammad Mohammadi, Ziyi Wu, Igor Gilitschenski
Abstract
Long-term temporal information is crucial for event-based perception tasks, as raw events only encode pixel brightness changes. Recent works show that when trained from scratch, recurrent models achieve better results than feedforward models in these tasks. However, when leveraging self-supervised pre-trained weights, feedforward models can outperform their recurrent counterparts. Current self-supervised learning (SSL) methods for event-based pre-training largely mimic RGB image-based approaches. They pre-train feedforward models on raw events within a short time interval, ignoring the temporal information of events. In this work, we introduce TESPEC, a self-supervised pre-training framework tailored for learning spatio-temporal information. TESPEC is well-suited for recurrent models, as it is the first framework to leverage long event sequences during pre-training. TESPEC employs the masked image modeling paradigm with a new reconstruction target. We design a novel method to accumulate events into pseudo grayscale videos containing high-level semantic information about the underlying scene, which is robust to sensor noise and reduces motion blur. Reconstructing this target thus requires the model to reason about long-term history of events. Extensive experiments demonstrate our state-of-the-art results in downstream tasks, including object detection, semantic segmentation, and monocular depth estimation. Project webpage: https://mhdmohammadi.github.io/TESPEC_webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Scaling Dense Event-Stream Pretraining from Visual Foundation ModelsZhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu et al.CVPR 2026 · 2 citations
- Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth EstimationDaikun Liu, Teng Wang, Changyin SunCVPR 2026
Builds on40
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse EventsLin Zhu, Ruonan Liu, Xiao Wang, Lizhi Wang et al.ACM MM 2025 · 1 citation
- Spatio-Temporal Recurrent Networks for Event-Based Optical Flow EstimationZiluo Ding, Rui Zhao, Jiyuan Zhang, Tianxiao Gao et al.AAAI 2022 · 76 citations
- DERD-Net: Learning Depth from Event-based Ray DensitiesDiego de Oliveira Hitzges, Suman Ghosh, Guillermo GallegoNeurIPS 2025 · 6 citations
- Self-Supervised Learning of Event-Based Optical Flow with Spiking Neural NetworksJesse J. Hagenaars, Federico Paredes-Vallés, Guido de CroonNeurIPS 2021 · 178 citations
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion ModelsCheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin et al.SIGGRAPH 2026
