Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Qiang Li, Shuai Liu, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang, Kai Huang, Yonghong Tian
摘要
Applying pretraining-finetuning paradigm to event cameras presents significant challenges due to the scarcity of largescale event datasets and the inherently sparse nature of event data, which increases the risk of overfitting during extensive pretraining. In this paper, we explore the transfer of pretrained image knowledge to the domain of event cameras to address this challenge. The key to our approach lies in adapting event data representations to align with image pretrained models while simultaneously integrating spatiotemporal information and mitigating data sparsity. To achieve this, we propose a lightweight SpatioTemporal information fusion Prompting (STP) method, which progressively fuses the spatiotemporal characteristics of event data through a dynamic perception module with multi-scale spatiotemporal receptive fields, enabling compatibility with image pretrained models. STP enhances event data representation by capturing local information within a large receptive field and performing global information exchange along the temporal dimension. This strategy effectively reduces sparse regions in event data while refining fine-grained details, all while preserving its inherent spatiotemporal structure. Our method significantly outperforms previous state-of-the-art approaches across classification, semantic segmentation, and optical flow estimation tasks. For instance, it achieves a top-1 accuracy of 68.87% (+4.04%) on N-ImageNet with only 1/10 of the pretraining parameters and 1/3 of the training epochs. Our code is available at https://github.com/Lqm26/STP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- EMatch: A Unified Framework for Event-Based Optical Flow and Stereo MatchingPengjie Zhang, Lin Zhu, Xiao Wang, Lizhi Wang 等ICCV 2025 · 被引用 2 次
- ESEG: Event-Based Segmentation Boosted by Explicit Edge-Semantic GuidanceYucheng Zhao, Gengyu Lyu, Ke Li, Zihao Wang 等AAAI 2025 · 被引用 8 次
- Event Camera Data Pre-trainingYan Yang, Liyuan Pan, Liu LiuICCV 2023 · 被引用 59 次
- Zero-Shot Event-Intensity Asymmetric Stereo via Visual Prompting from Image DomainHanyue Lou, Jinxiu (Sherry) Liang, Minggui Teng, Bin Fan 等NeurIPS 2024 · 被引用 13 次
- Revealing Latent Information: A Physics-inspired Self-supervised Pre-training Framework for Noisy and Sparse EventsLin Zhu, Ruonan Liu, Xiao Wang, Lizhi Wang 等ACM MM 2025 · 被引用 1 次
