Reformulating Zero-shot Action Recognition for Multi-label Actions
Alec Kerrigan, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah
Abstract
The goal of zero-shot action recognition (ZSAR) is to classify action classes which were not previously seen during training. Traditionally, this is achieved by training a network to map, or regress, visual inputs to a semantic space where a nearest neighbor classifier is used to select the closest target class. We argue that this approach is sub-optimal due to the use of nearest neighbor on static semantic space and is ineffective when faced with multi-label videos -where two semantically distinct co-occurring action categories cannot be predicted with high confidence. To overcome these limitations, we propose a ZSAR framework which does not rely on nearest neighbor classification, but rather consists of a pairwise scoring function. Given a video and a set of action classes, our method predicts a set of confidence scores for each class independently. This allows for the prediction of several semantically distinct classes within one video input. Our evaluations show that our method not only achieves strong performance on three single-label action classification datasets (UCF-101, HMDB, and RareAct), but also outperforms previous ZSAR approaches on a challenging multi-label dataset (AVA) and a real-world surprise activity detection dataset (MEVA).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
- Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationZiwei Xu, Yogesh S. Rawat, Yongkang Wong, Mohan S. Kankanhalli et al.NeurIPS 2022 · 18 citations
- Uncovering the Hidden Dynamics of Video Self-supervised Learning under Distribution ShiftsPritam Sarkar, Ahmad Beirami, Ali EtemadNeurIPS 2023 · 8 citations
- Punching Bag vs. Punching Person: Motion Transferability in VideosRaiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran et al.ICCV 2025 · 1 citation
Builds on4
- A Shared Multi-Attention Framework for Multi-Label Zero-Shot LearningDat Huynh, Ehsan ElhamifarCVPR 2020
- Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic ApplicationsBiagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona et al.CVPR 2020
- Domain-Aware Visual Bias Eliminating for Generalized Zero-Shot LearningShaobo Min, Hantao Yao, Hongtao Xie, Chaoqun Wang et al.CVPR 2020
- Self-Supervised Domain-Aware Generative Network for Generalized Zero-Shot LearningJiamin Wu, Tianzhu Zhang, Zheng-Jun Zha, Jiebo Luo et al.CVPR 2020
Related papers
- Crossmodal Representation Learning for Zero-shot Action RecognitionChung-Ching Lin, Kevin Lin, Lijuan Wang, Zicheng Liu et al.CVPR 2022 · 39 citations
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
- Elaborative Rehearsal for Zero-shot Action RecognitionShizhe Chen, Dong HuangICCV 2021 · 114 citations
- Few-Shot Transformation of Common Actions Into Time and SpacePengwan Yang, Pascal Mettes, Cees G. M. SnoekCVPR 2021
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang et al.ACM MM 2022 · 11 citations
