Class Semantics-based Attention for Action Detection
Deepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li, Peng Dai, Juwei Lu
摘要
Action localization networks are often structured as a feature encoder sub-network and a localization sub-network, where the feature encoder learns to transform an input video to features that are useful for the localization sub-network to generate reliable action proposals. While some of the encoded features may be more useful for generating action proposals, prior action localization approaches do not include any attention mechanism that enables the localization sub-network to attend more to the more important features. In this paper, we propose a novel attention mechanism, the Class Semantics-based Attention (CSA), that learns from the temporal distribution of semantics of action classes present in an input video to find the importance scores of the encoded features, which are used to provide attention to the more useful encoded features. We demonstrate on two popular action detection datasets that incorporating our novel attention mechanism provides considerable performance gains on competitive action detection models (e.g., around 6.2% improvement over BMN action detection baseline to obtain 47.5% mAP on the THUMOS-14 dataset), and a new state-of-the-art of 36.25% mAP on the ActivityNet v1.3 dataset. Further, the CSA localization model family which includes BMN-CSA, was part of the second-placed submission at the 2021 ActivityNet action localization challenge. Our attention mechanism outperforms prior self-attention modules such as the squeeze-and-excitation in action detection task. We also observe that our attention mechanism is complementary to such self-attention modules in that performance improvements are seen when both are used together.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action LocalizationJunyu Gao, Mengyuan Chen, Changsheng XuCVPR 2022 · 被引用 87 次
- An Empirical Study of End-to-End Temporal Action DetectionXiaolong Liu, Song Bai, Xiang BaiCVPR 2022 · 被引用 72 次
- TORE: Token Reduction for Efficient Human Mesh Recovery with TransformerZhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao 等ICCV 2023 · 被引用 56 次
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 被引用 50 次
- Action Sensitivity Learning for Temporal Action LocalizationJiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng 等ICCV 2023 · 被引用 44 次
它引用的顶会 Paper10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
相关 Paper
- A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationZan Gao, Xinglei Cui, Yibo Zhao, Tao Zhuo 等ACM MM 2023 · 被引用 2 次
- Deep Concept-wise Temporal Convolutional Networks for Action LocalizationXin Li, Tianwei Lin, Xiao Liu, Wangmeng Zuo 等ACM MM 2020 · 被引用 27 次
- Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention ModelingGuiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang 等ICCV 2023 · 被引用 7 次
- MTSN: Multiscale Temporal Similarity Network for Temporal Action LocalizationXiaodong Jin, Taiping ZhangACM MM 2023 · 被引用 3 次
- Weakly-Supervised Action Localization by Generative Attention ModelingBaifeng Shi, Qi Dai, Yadong Mu, Jingdong WangCVPR 2020
