EZSR: Event-based Zero-Shot Recognition
Yan Yang, Liyuan Pan, Dongxu Li, Liu Liu
Abstract
This paper studies zero-shot object recognition using event camera data. Guided by CLIP, which is pre-trained on RGB images, existing approaches achieve zero-shot object recognition by optimizing embedding similarities between event data and RGB images respectively encoded by an event encoder and the CLIP image encoder. Alternatively, several methods learn RGB frame reconstructions from event data for the CLIP image encoder. However, they often result in suboptimal zero-shot performance. This study develops an event encoder without relying on additional reconstruction networks. We theoretically analyze the performance bottlenecks of previous approaches: the embedding optimization objectives are prone to suffer from the spatial sparsity of event data, causing semantic misalignments between the learned event embedding space and the CLIP text embedding space. To mitigate the issue, we explore a scalar-wise modulation strategy. Furthermore, to scale up the number of events and RGB data pairs for training, we also study a pipeline for synthesizing event data from static RGB images in mass. Experimentally, we demonstrate an attractive scaling property in the number of parameters and synthesized data. We achieve superior zero-shot object recognition performance on extensive standard benchmark datasets, even compared with past supervised learning approaches. For example, our model with a ViT/B-16 backbone achieves 47.84% zero-shot accuracy on the N-ImageNet dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Scaling Dense Event-Stream Pretraining from Visual Foundation ModelsZhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu et al.CVPR 2026 · 2 citations
- Seeing Motion Through Polarity for Event-based Action RecognitionMeiqi Cao, Jiachao Zhang, Xin Jiang, Rui Yan et al.CVPR 2026
- RE-VLM: Event-Augmented Vision-Language Model for Scene UnderstandingHanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin et al.CVPR 2026
- DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera ConditionsJiaqi Chen, Qinfu Xu, Liyuan PanCVPR 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from EventsHoonhee Cho, Hyeonseong Kim, Yujeong Chae, Kuk-Jin YoonICCV 2023 · 38 citations
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu et al.ICML 2023 · 67 citations
- Event Camera Data Pre-trainingYan Yang, Liyuan Pan, Liu LiuICCV 2023 · 59 citations
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
