EZSR: Event-based Zero-Shot Recognition
Yan Yang, Liyuan Pan, Dongxu Li, Liu Liu
摘要
This paper studies zero-shot object recognition using event camera data. Guided by CLIP, which is pre-trained on RGB images, existing approaches achieve zero-shot object recognition by optimizing embedding similarities between event data and RGB images respectively encoded by an event encoder and the CLIP image encoder. Alternatively, several methods learn RGB frame reconstructions from event data for the CLIP image encoder. However, they often result in suboptimal zero-shot performance. This study develops an event encoder without relying on additional reconstruction networks. We theoretically analyze the performance bottlenecks of previous approaches: the embedding optimization objectives are prone to suffer from the spatial sparsity of event data, causing semantic misalignments between the learned event embedding space and the CLIP text embedding space. To mitigate the issue, we explore a scalar-wise modulation strategy. Furthermore, to scale up the number of events and RGB data pairs for training, we also study a pipeline for synthesizing event data from static RGB images in mass. Experimentally, we demonstrate an attractive scaling property in the number of parameters and synthesized data. We achieve superior zero-shot object recognition performance on extensive standard benchmark datasets, even compared with past supervised learning approaches. For example, our model with a ViT/B-16 backbone achieves 47.84% zero-shot accuracy on the N-ImageNet dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Scaling Dense Event-Stream Pretraining from Visual Foundation ModelsZhiwen Chen, Junhui Hou, Zhiyu Zhu, Jinjian Wu 等CVPR 2026 · 被引用 2 次
- Seeing Motion Through Polarity for Event-based Action RecognitionMeiqi Cao, Jiachao Zhang, Xin Jiang, Rui Yan 等CVPR 2026
- RE-VLM: Event-Augmented Vision-Language Model for Scene UnderstandingHanqing Liu, Mingjie Liu, Luoping Cui, Endian Lin 等CVPR 2026
- DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera ConditionsJiaqi Chen, Qinfu Xu, Liyuan PanCVPR 2026
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from EventsHoonhee Cho, Hyeonseong Kim, Yujeong Chae, Kuk-Jin YoonICCV 2023 · 被引用 38 次
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 等ICML 2023 · 被引用 67 次
- Event Camera Data Pre-trainingYan Yang, Liyuan Pan, Liu LiuICCV 2023 · 被引用 59 次
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang 等CVPR 2023
