Fine-grained Temporal Contrastive Learning for Weakly-supervised Temporal Action Localization
Junyu Gao, Mengyuan Chen, Changsheng Xu
摘要
We target at the task of weakly-supervised action localization (WSAL), where only video-level action labels are available during model training. Despite the recent progress, existing methods mainly embrace a localization-by-classification paradigm and overlook the fruitful fine-grained temporal distinctions between video sequences, thus suffering from severe ambiguity in classification learning and classification-to-localization adaption. This paper argues that learning by contextually comparing sequence-to-sequence distinctions offers an essential inductive bias in WSAL and helps identify coherent action instances. Specifically, under a differentiable dynamic programming formulation, two complementary contrastive objectives are designed, including Fine-grained Sequence Distance (FSD) contrasting and Longest Common Subsequence (LCS) contrasting, where the first one considers the relations of various action/background proposals by using match, insert, and delete operators and the second one mines the longest common subsequences between two videos. Both contrasting modules can enhance each other and jointly enjoy the merits of discriminative action-background separation and alleviated task gap between classification and localization. Extensive experiments show that our method achieves state-of-the-art performance on two popular benchmarks. Our code is available at https://github.com/MengyuanChen21/CVPR2022-FTCL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang 等CVPR 2024 · 被引用 31 次
- Forcing the Whole Video as Background: An Adversarial Learning Strategy for Weakly Temporal Action LocalizationZiqiang Li, Yongxin Ge, Jiaruo Yu, Zhongming ChenACM MM 2022 · 被引用 24 次
- Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D SpaceJinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 被引用 18 次
- Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based ApproachQinying Liu, Zilei Wang, Shenghai Rong, Junjie Li 等ICCV 2023 · 被引用 18 次
- DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationXiaojun Tang, Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang 等ICCV 2023 · 被引用 16 次
它引用的顶会 Paper28
- Background Suppression Network for Weakly-Supervised Temporal Action LocalizationPilhyeon Lee, Youngjung Uh, Hyeran ByunAAAI 2020 · 被引用 234 次
- Relaxed Transformer Decoders for Direct Action Proposal GenerationJing Tan, Jiaqi Tang, Limin Wang, Gangshan WuICCV 2021 · 被引用 220 次
- Weakly-Supervised Action Localization With Background ModelingPhuc Xuan Nguyen, Deva Ramanan, Charless C. FowlkesICCV 2019 · 被引用 176 次
- 3C-Net: Category Count and Center Loss for Weakly-Supervised Action LocalizationSanath Narayan, Hisham Cholakkal, Fahad Shahbaz Khan, Ling ShaoICCV 2019 · 被引用 174 次
- A Hybrid Attention Mechanism for Weakly-Supervised Temporal Action LocalizationAshraful Islam, Chengjiang Long, Richard J. RadkeAAAI 2021 · 被引用 145 次
相关 Paper
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 被引用 11 次
- CoLA: Weakly-Supervised Temporal Action Localization With Snippet Contrastive LearningCan Zhang, Meng Cao, Dongming Yang, Jie Chen 等CVPR 2021
- Actionness Inconsistency-Guided Contrastive Learning for Weakly-Supervised Temporal Action LocalizationZhilin Li, Zilei Wang, Qinying LiuAAAI 2023 · 被引用 12 次
- Exploring Denoised Cross-video Contrast for Weakly-supervised Temporal Action LocalizationJingjing Li, Tianyu Yang, Wei Ji, Jue Wang 等CVPR 2022 · 被引用 57 次
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 被引用 34 次
