Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic Applications
Biagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona, Krzysztof Chalupka
摘要
Trained on large datasets, deep learning (DL) can accurately classify videos into hundreds of diverse classes. However, video data is expensive to annotate. Zero-shot learning (ZSL) proposes one solution to this problem. ZSL trains a model once, and generalizes to new tasks whose classes are not present in the training dataset. We propose the first end-to-end algorithm for ZSL in video classification. Our training procedure builds on insights from recent video classification literature and uses a trainable 3D CNN to learn the visual features. This is in contrast to previous video ZSL methods, which use pretrained feature extractors. We also extend the current benchmarking paradigm: Previous techniques aim to make the test task unknown at training time but fall short of this goal. We encourage domain shift across training and test data and disallow tailoring a ZSL model to a specific test dataset. We outperform the state-of-the-art by a wide margin. Our code, evaluation procedure and model weights are available at github.com/bbrattoli/ZeroShotVideoClassification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the BackboneShraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin 等ICCV 2023 · 被引用 152 次
- Revisiting Classifier: Transferring Vision-Language Models for Video RecognitionWenhao Wu, Zhun Sun, Wanli OuyangAAAI 2023 · 被引用 141 次
- Elaborative Rehearsal for Zero-shot Action RecognitionShizhe Chen, Dong HuangICCV 2021 · 被引用 114 次
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne 等ICCV 2021 · 被引用 98 次
- Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight OptimizationZejia Weng, Xitong Yang, Ang Li, Zuxuan Wu 等ICML 2023 · 被引用 67 次
它引用的顶会 Paper5
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Revisiting Training Strategies and Generalization Performance in Deep Metric LearningKarsten Roth, Timo Milbich, Samarth Sinha, Prateek Gupta 等ICML 2020 · 被引用 187 次
- MIC: Mining Interclass Characteristics for Improved Metric LearningBiagio Brattoli, Karsten Roth, Björn OmmerICCV 2019 · 被引用 100 次
- Action Recognition With Spatial-Temporal Discriminative Filter BanksBrais Martínez, Davide Modolo, Yuanjun Xiong, Joseph TigheICCV 2019 · 被引用 70 次
- Transferring Dense Pose to Proximal Animal ClassesArtsiom Sanakoyeu, Vasil Khalidov, Maureen S. McCarthy, Andrea Vedaldi 等CVPR 2020
相关 Paper
- Generalized Zero-Shot Video Classification via Generative Adversarial NetworksMingyao Hong, Guorong Li, Xinfeng Zhang, Qingming HuangACM MM 2020 · 被引用 13 次
- Rethinking Zero-Shot Learning: A Conditional Visual Classification PerspectiveKai Li, Martin Renqiang Min, Yun FuICCV 2019 · 被引用 151 次
- A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningPeirong Ma, Xiao HuAAAI 2020 · 被引用 43 次
- Zero-shot Video Classification with Appropriate Web and Task Knowledge TransferJunbao Zhuo, Yan Zhu, Shuhao Cui, Shuhui Wang 等ACM MM 2022 · 被引用 11 次
- Audiovisual Generalised Zero-shot Learning with Cross-modal Attention and LanguageOtniel-Bogdan Mercea, Lukas Riesch, A. Sophia Koepke, Zeynep AkataCVPR 2022 · 被引用 54 次
