Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention Modeling
Guiqin Wang, Peng Zhao, Cong Zhao, Shusen Yang, Jie Cheng, Luziwei Leng, Jianxing Liao, Qinghai Guo
摘要
Weakly-supervised action localization aims to recognize and localize action instancese in untrimmed videos with only video-level labels. Most existing models rely on multiple instance learning(MIL), where the predictions of unlabeled instances are supervised by classifying labeled bags. The MIL-based methods are relatively well studied with cogent performance achieved on classification but not on localization. Generally, they locate temporal regions by the video-level classification but overlook the temporal variations of feature semantics. To address this problem, we propose a novel attention-based hierarchically-structured latent model to learn the temporal variations of feature semantics. Specifically, our model entails two components, the first is an unsupervised change-points detection module that detects change-points by learning the latent representations of video features in a temporal hierarchy based on their rates of change, and the second is an attention-based classification model that selects the change-points of the foreground as the boundaries. To evaluate the effectiveness of our model, we conduct extensive experiments on two benchmark datasets, THUMOS-14 and ActivityNet-v1.3. The experiments show that our method outperforms current state-of-the-art methods, and even achieves comparable performance with fully-supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 被引用 11 次
- Action-Agnostic Point-Level Supervision for Temporal Action DetectionShuhei M. Yoshida, Takashi Shibata, Makoto Terao, Takayuki Okatani 等AAAI 2025 · 被引用 6 次
- Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language ModelsQuan Zhang, Jinwei Fang, Rui Yuan, Xi Tang 等CVPR 2025
它引用的顶会 Paper23
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Graph Convolutional Networks for Temporal Action LocalizationRunhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan 等ICCV 2019 · 被引用 536 次
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 被引用 234 次
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 被引用 176 次
- 3C-Net: Category Count and Center Loss for Weakly-Supervised Action LocalizationSanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, Ling ShaoICCV 2019 · 被引用 174 次
相关 Paper
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
- Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsYu Wang, Yadong Li, Hongbin WangCVPR 2023
- A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationAshraful Islam, Chengjiang Long, Richard J. RadkeAAAI 2021 · 被引用 145 次
- Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed VideosXiaoyu Zhang, Haichao Shi, Changsheng Li, Peng LiAAAI 2020 · 被引用 37 次
- Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action LocalizationHuan Ren, Wenfei Yang, Tianzhu Zhang, Yongdong ZhangCVPR 2023
