ZSTAD: Zero-Shot Temporal Activity Detection
Lingling Zhang, Xiaojun Chang, Jun Liu, Minnan Luo, Sen Wang, Zongyuan Ge, Alexander G. Hauptmann
摘要
An integral part of video analysis and surveillance is temporal activity detection, which means to simultaneously recognize and localize activities in long untrimmed videos. Currently, the most effective methods of temporal activity detection are based on deep learning, and they typically perform very well with large scale annotated videos for training. However, these methods are limited in real applications due to the unavailable videos about certain activity classes and the time-consuming data annotation. To solve this challenging problem, we propose a novel task setting called zero-shot temporal activity detection (ZSTAD), where activities that have never been seen in training can still be detected. We design an end-to-end deep network based on R-C3D as the architecture for this solution. The proposed network is optimized with an innovative loss function that considers the embeddings of activity labels and their superclasses while learning the common semantics of seen and unseen activities. Experiments on both the THUMOS'14 and the Charades datasets show promising performance in terms of detecting unseen activities.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pseudo-Q: Generating Pseudo Language Queries for Visual GroundingHaojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song 等CVPR 2022 · 被引用 60 次
- Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action DetectionYearang Lee, Ho-Joong Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 12 次
相关 Paper
- Weakly-Guided Self-Supervised Pretraining for Temporal Activity DetectionKumara Kahatapitiya, Zhou Ren, Haoxiang Li, Zhenyu Wu 等AAAI 2023 · 被引用 7 次
- ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionYongqi An, Xu Zhao, Tao Yu, Haiyun Gu 等CVPR 2023
- Three Birds with One Stone: Multi-Task Temporal Action Detection via Recycling Temporal AnnotationsZhihui Li, Lina YaoCVPR 2021
- Rethinking Zero-Shot Video Classification: End-to-End Training for Realistic ApplicationsBiagio Brattoli, Joseph Tighe, Fedor Zhdanov, Pietro Perona 等CVPR 2020
- WOAD: Weakly Supervised Online Action Detection in Untrimmed VideosMingfei Gao, Yingbo Zhou, Ran Xu, Richard Socher 等CVPR 2021
