Modeling Multi-Label Action Dependencies for Temporal Action Localization
Praveen Tirupattur, Kevin Duarte, Yogesh S. Rawat, Mubarak Shah
摘要
Real-world videos contain many complex actions with inherent relationships between action classes. In this work, we propose an attention-based architecture that models these action relationships for the task of temporal action localization in untrimmed videos. As opposed to previous works that leverage video-level co-occurrence of actions, we distinguish the relationships between actions that occur at the same time-step and actions that occur at different time-steps (i.e. those which precede or follow each other). We define these distinct relationships as action dependencies. We propose to improve action localization performance by modeling these action dependencies in a novel attention-based Multi-Label Action Dependency (MLAD) layer. The MLAD layer consists of two branches: a Cooccurrence Dependency Branch and a Temporal Dependency Branch to model co-occurrence action dependencies and temporal action dependencies, respectively. We observe that existing metrics used for multi-label classification do not explicitly measure how well action dependencies are modeled, therefore, we propose novel metrics that consider both co-occurrence and temporal dependencies between action classes. Through empirical evaluation and extensive analysis, we show improved performance over state-of-the-art methods on multi-label action localization benchmarks (MultiTHUMOS and Charades) in terms of f-mAP and our proposed metric. Code is publicly available at https://github.com/ptirupat/MLAD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Exploring Cross-Video and Cross-Modality Signals for Weakly-Supervised Audio-Visual Video ParsingYan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin 等NeurIPS 2021 · 被引用 94 次
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo 等CVPR 2022 · 被引用 93 次
- Action Sensitivity Learning for Temporal Action LocalizationJiayi Shao, Xiaohan Wang, Ruijie Quan, Junjun Zheng 等ICCV 2023 · 被引用 44 次
- Learning to Refactor Action and Co-occurrence Features for Temporal Action LocalizationKun Xia, Le Wang, Sanping Zhou, Nanning Zheng 等CVPR 2022 · 被引用 43 次
- PointTAD: Multi-Label Temporal Action Detection with Learnable Query PointsJing Tan, Xiaotong Zhao, Xintian Shi, Bin Kang 等NeurIPS 2022 · 被引用 41 次
它引用的顶会 Paper7
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan 等ICCV 2019 · 被引用 536 次
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 被引用 298 次
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long 等AAAI 2020 · 被引用 221 次
- Differentiable Grammars for VideosA. J. Piergiovanni, Anelia Angelova, Michael S. RyooAAAI 2020 · 被引用 6 次
相关 Paper
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li 等AAAI 2025 · 被引用 3 次
- Learning Temporal Co-Attention Models for Unsupervised Video Action LocalizationGuoqiang Gong, Xinghan Wang, Yadong Mu, Qi TianCVPR 2020
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Class Semantics-based Attention for Action DetectionDeepak Sridhar, Niamul Quader, Srikanth Muralidharan, Yaoxin Li 等ICCV 2021 · 被引用 77 次
- Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention ModelingGuiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang 等ICCV 2023 · 被引用 7 次
