Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic Applications
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, Krzysztof Chalupka
Abstract
Trained on large datasets, deep learning (DL) can accurately classify videos into hundreds of diverse classes. However, video data is expensive to annotate. Zero-shot learning (ZSL) proposes one solution to this problem. ZSL trains a model once, and generalizes to new tasks whose classes are not present in the training dataset. We propose the first end-to-end algorithm for ZSL in video classification. Our training procedure builds on insights from recent video classification literature and uses a trainable 3D CNN to learn the visual features. This is in contrast to previous video ZSL methods, which use pretrained feature extractors. We also extend the current benchmarking paradigm: Previous techniques aim to make the test task unknown at training time but fall short of this goal. We encourage domain shift across training and test data and disallow tailoring a ZSL model to a specific test dataset. We outperform the state-of-the-art by a wide margin. Our code, evaluation procedure and model weights are available at github.com/bbrattoli/ZeroShotVideoClassification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers41
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin et al.ICCV 2023 · 152 citations
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 141 citations
- Elaborative Rehearsal for Zero-shot Action RecognitionShizhe Chen, Dong HuangICCV 2021 · 114 citations
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne et al.ICCV 2021 · 98 citations
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu et al.ICML 2023 · 67 citations
Builds on5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Revisiting Training Strategies and Generalization Performance in Deep Metric LearningKarsten Roth, Timo Milbich, Samarth Sinha, Prateek Gupta et al.ICML 2020 · 187 citations
- MIC: Mining Interclass Characteristics for Improved Metric LearningBiagio Brattoli, Karsten Roth, Björn OmmerICCV 2019 · 100 citations
- Action Recognition With Spatial-Temporal Discriminative Filter BanksBrais Martínez, Davide Modolo, Yuanjun Xiong, Joseph TigheICCV 2019 · 70 citations
- Transferring Dense Pose to Proximal Animal ClassesArtsiom Sanakoyeu, Vasil Khalidov, Maureen S. McCarthy, Andrea Vedaldi et al.CVPR 2020
Related papers
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 13 citations
- Rethinking Zero-Shot Learning: A Conditional Visual Classification PerspectiveKai Li, Martin Renqiang Min, Yun FuICCV 2019 · 151 citations
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 43 citations
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang et al.ACM MM 2022 · 11 citations
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 54 citations
