Event Camera Data Pre-training
Yan Yang, Liyuan Pan, Liu Liu
Abstract
This paper proposes a pre-trained neural network for handling event camera data. Our model is a self-supervised learning framework, and uses paired event camera data and natural RGB images for training. Our method contains three modules connected in a sequence: i) a family of event data augmentations, generating meaningful event images for self-supervised training; ii) a conditional masking strategy to sample informative event patches from event images, encouraging our model to capture the spatial layout of a scene and accelerating training; iii) a contrastive learning approach, enforcing the similarity of embeddings between matching event images, and between paired event and RGB images. An embedding projection loss is proposed to avoid the model collapse when enforcing the event image embedding similarities. A probability distribution alignment loss is proposed to encourage the event image to be consistent with its paired RGB image in the feature space. Transfer learning performance on downstream tasks shows the superiority of our method over state-of-the-art methods. For example, we achieve top-1 accuracy at 64.83% on the N-ImageNet dataset. Our code is available at https://github.com/Yan98/Event-Camera-Data-Pre-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c862c3ac-5f14-47c2-af1d-d5e9776f547bCited by top-tier papers21
- Language-driven All-in-one Adverse Weather RemovalHao Yang, Liyuan Pan, Yan Yang, Wei LiangCVPR 2024 · 28 citations
- FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational FrequenciesDongyue Lu, Lingdong Kong, Gim Hee Lee, Camille Simon Chane et al.NeurIPS 2025 · 13 citations
- Efficient Meshflow and Optical Flow Estimation from Event CamerasXinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin et al.CVPR 2024 · 10 citations
- LEOD: Label-Efficient Object Detection for Event CamerasZiyi Wu, Mathias Gehrig, Qing Lyu, Xudong Liu et al.CVPR 2024 · 10 citations
- LDP: Language-driven Dual-Pixel Image Defocus Deblurring NetworkHao Yang, Liyuan Pan, Yan Yang, Richard I. Hartley et al.CVPR 2024 · 9 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse EventsLin Zhu, Ruonan Liu, Xiao Wang, Lizhi Wang et al.ACM MM 2025 · 1 citation
- CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training FrameworkWentao Wu, Xiao Wang, Chenglong Li, Bo Jiang et al.ACM MM 2025 · 2 citations
- Unsupervised Domain Adaptation for Training Event-Based Networks Using Contrastive Learning and Uncorrelated ConditioningDayuan Jian, Mohammad RostamiICCV 2023 · 22 citations
- EZSR: Event-based Zero-Shot RecognitionYan Yang, Liyuan Pan, Dongxu Li, Liu LiuCVPR 2025
- N-ImageNet: Towards Robust, Fine-Grained Object Recognition with Event CamerasJunho Kim, Jaehyeok Bae, Gangin Park, Dongsu Zhang et al.ICCV 2021 · 127 citations
