Learning Disentangled Classification and Localization Representations for Temporal Action Localization
Zixin Zhu, Le Wang, Wei Tang, Ziyi Liu, Nanning Zheng, Gang Hua
摘要
A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., take-offs" rather than run-ups" in distinguishing high jump" and long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Jae-Pil HeoICCV 2023 · 被引用 33 次
- Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action RecognitionHaochen Chang, Pengfei Ren, Haoyang Zhang, Liang Xie 等ICCV 2025 · 被引用 8 次
- Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormerZiyi Liu, Yangcen LiuCVPR 2025
它引用的顶会 Paper18
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan 等ICCV 2019 · 被引用 536 次
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai 等AAAI 2020 · 被引用 226 次
- Progressive Boundary Refinement Network for Temporal Action DetectionQinying Liu, Zilei WangAAAI 2020 · 被引用 156 次
- BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal GenerationHaisheng Su, Weihao Gan, Wei Wu, Yu Qiao 等AAAI 2021 · 被引用 143 次
相关 Paper
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li 等AAAI 2025 · 被引用 3 次
- Enriching Local and Global Contexts for Temporal Action LocalizationZixin Zhu, Wei Tang, Le Wang, Nanning Zheng 等ICCV 2021 · 被引用 134 次
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan 等AAAI 2021 · 被引用 28 次
- Divide and Conquer for Single-frame Temporal Action LocalizationChen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 等ICCV 2021 · 被引用 46 次
- Weakly-Supervised Action Localization by Generative Attention ModelingBaifeng Shi, Qi Dai, Yadong Mu, Jingdong WangCVPR 2020
