Forcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action Localization
Ziqiang Li, Yongxin Ge, Jiaruo Yu, Zhongming Chen
Abstract
With video-level labels, weakly supervised temporal action localization (WTAL) applies a localization-by-classification paradigm to detect and classify the action in untrimmed videos. Due to the characteristic of classification, class-specific background snippets are inevitably mis-activated to improve the discriminability of the classifier in WTAL. To alleviate the disturbance of background, existing methods try to enlarge the discrepancy between action and background through modeling background snippets with pseudo-snippet-level annotations, which largely rely on artificial hypotheticals. Distinct from the previous works, we present an adversarial learning strategy to break the limitation of mining pseudo background snippets. Concretely, the background classification loss forces the whole video to be regarded as the background by a background gradient reinforcement strategy, confusing the recognition model. Reversely, the foreground(action) loss guides the model to focus on action snippets under such conditions. As a result, competition between the two classification losses drives the model to boost its ability for action modeling. Simultaneously, a novel temporal enhancement network is designed to facilitate the model to construct temporal relation of affinity snippets based on the proposed strategy, for further improving the performance of action localization. Finally, extensive experiments conducted on THUMOS14 and ActivityNet1.2 demonstrate the effectiveness of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationXiaojun Tang, Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang et al.ICCV 2023 · 16 citations
- Boosting Positive Segments for Weakly-Supervised Audio-Visual Video ParsingKranthi Kumar Rachavarapu, A. N. RajagopalanICCV 2023 · 13 citations
- Improving Weakly Supervised Temporal Action Localization by Bridging Train-Test Gap in Pseudo LabelsJingqiu Zhou, Linjiang Huang, Liang Wang, Si Liu et al.CVPR 2023
Builds on21
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan et al.ICCV 2019 · 536 citations
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 234 citations
- Fast Learning of Temporal Action Proposal via Dense Boundary GeneratorChuming Lin, Jian Li, Yabiao Wang, Ying Tai et al.AAAI 2020 · 226 citations
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 176 citations
Related papers
- PivoTAL: Prior-Driven Supervision for Weakly-Supervised Temporal Action LocalizationMamshad Nayeem Rizve, Gaurav Mittal, Ye Yu, Matthew Hall et al.CVPR 2023
- ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationZiyi Liu, Le Wang, Qilin Zhang, Wei Tang et al.AAAI 2021 · 83 citations
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan et al.AAAI 2021 · 28 citations
- Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based ApproachQinying Liu, Zilei Wang, Shenghai Rong, Junjie Li et al.ICCV 2023 · 18 citations
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng et al.CVPR 2022 · 104 citations
