Learning Disentangled Classification and Localization Representations for Temporal Action Localization
Zixin Zhu, Le Wang, Wei Tang, Ziyi Liu, Nanning Zheng, Gang Hua
Abstract
A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., take-offs" rather than run-ups" in distinguishing high jump" and long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Self-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Jae-Pil HeoICCV 2023 · 33 citations
- Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action RecognitionHaochen Chang, Pengfei Ren, Haoyang Zhang, Liang Xie et al.ICCV 2025 · 8 citations
- Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormerZiyi Liu, Yangcen LiuCVPR 2025
Builds on18
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
- Progressive Boundary Refinement Network for Temporal Action DetectionQinying Liu, Zilei WangAAAI 2020 · 156 citations
- BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal GenerationHaisheng Su, Weihao Gan, Wei Wu, Yu Qiao et al.AAAI 2021 · 143 citations
Related papers
- Temporal Action Localization with Cross Layer Task Decoupling and RefinementQiang Li, Di Liu, Jun Kong, Sen Li et al.AAAI 2025 · 3 citations
- Enriching Local and Global Contexts for Temporal Action LocalizationZixin Zhu, Wei Tang, Le Wang, Nanning Zheng et al.ICCV 2021 · 134 citations
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
- Divide and Conquer for Single-frame Temporal Action LocalizationChen Ju, Peisen Zhao, Siheng Chen, Ya Zhang et al.ICCV 2021 · 46 citations
- Weakly-Supervised Action Localization by Generative Attention ModelingBaifeng Shi, Qi Dai, Yadong Mu, Jingdong WangCVPR 2020
