Elaborative Rehearsal for Zero-shot Action Recognition
Shizhe Chen, Dong Huang
Abstract
The growing number of action classes has posed a new challenge for video understanding, making Zero-Shot Action Recognition (ZSAR) a thriving direction. The ZSAR task aims to recognize target (unseen) actions without training examples by leveraging semantic representations to bridge seen and unseen actions. However, due to the complexity and diversity of actions, it remains challenging to semantically represent action classes and transfer knowledge from seen data. In this work, we propose an ER-enhanced ZSAR model inspired by an effective human memory technique Elaborative Rehearsal (ER), which involves elaborating a new concept and relating it to known concepts. Specifically, we expand each action class as an Elaborative Description (ED) sentence, which is more discriminative than a class name and less costly than manual-defined attributes. Besides directly aligning class semantics with videos, we incorporate objects from the video as Elaborative Concepts (EC) to improve video semantics and generalization from seen actions to unseen actions. Our ER-enhanced ZSAR model achieves state-of-the-art results on three existing benchmarks. Moreover, we propose a new ZSAR evaluation protocol on the Kinetics dataset to overcome limitations of current benchmarks and first compare with few-shot learning baselines on this more realistic setting. Our codes and collected EDs are released at https://github. com/DeLightCMU/ElaborativeRehearsal .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 141 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
- VideoPrism: A Foundational Visual Encoder for Video UnderstandingLong Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou et al.ICML 2024 · 91 citations
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu et al.ICML 2023 · 67 citations
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song et al.CVPR 2022 · 60 citations
Builds on3
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Fine-Grained Generalized Zero-Shot Learning via Dense Attribute-Based AttentionDat Huynh, Ehsan ElhamifarCVPR 2020
- Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic ApplicationsBiagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona et al.CVPR 2020
Related papers
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang et al.ACM MM 2022 · 11 citations
- Crossmodal Representation Learning for Zero-shot Action RecognitionChung-Ching Lin, Kevin Lin, Lijuan Wang, Zicheng Liu et al.CVPR 2022 · 39 citations
- Attend and Enrich: Enhanced Visual Prompt for Zero-Shot LearningMan Liu, Huihui Bai, Feng Li, Chunjie Zhang et al.AAAI 2025 · 3 citations
- Reformulating Zero-shot Action Recognition for Multi-label ActionsAlec Kerrigan, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahNeurIPS 2021 · 23 citations
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
