Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action Localization
Chen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao, Ya Zhang, Jianlong Chang, Qi Tian, Yanfeng Wang
摘要
Weakly-supervised temporal action localization (WTAL) learns to detect and classify action instances with only category labels. Most methods widely adopt the off-the-shelf Classification-Based Pre-training (CBP) to generate video features for action localization. However, the different optimization objectives between classification and localization, make temporally localized results suffer from the serious incomplete issue. To tackle this issue without additional annotations, this paper considers to distill free action knowledge from Vision-Language Pre-training (VLP), as we surprisingly observe that the localization results of vanilla VLP have an over-complete issue, which is just complementary to the CBP results. To fuse such complementarity, we propose a novel distillation-collaboration framework with two branches acting as CBP and VLP respectively. The framework is optimized through a dual-branch alternate training strategy. Specifically, during the B step, we distill the confident background pseudo-labels from the CBP branch; while during the F step, the confident foreground pseudo-labels are distilled from the VLP branch. As a result, the dualbranch complementarity is effectively fused to promote one strong alliance. Extensive experiments and ablation studies on THUMOS14 and ActivityNet1.2 reveal that our method significantly outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Open-Vocabulary Semantic Segmentation via Attribute Decomposition-AggregationChaofan Ma, Yuhuan Yang, Chen Ju, Fei Zhang 等NeurIPS 2023 · 被引用 40 次
- X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferLinglin Jing, Ying Xue, Xu Yan, Chaoda Zheng 等AAAI 2024 · 被引用 14 次
- Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise CorrectionQuan Zhang, Yuxin Qi, Xi Tang, Rui Yuan 等AAAI 2025 · 被引用 11 次
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 被引用 11 次
- FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced PerformanceHaicheng Wang, Zhemeng Yu, Gabriele Spadaro, Chen Ju 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
相关 Paper
- Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextZiyi Liu, Le Wang, Wei Tang, Junsong Yuan 等AAAI 2021 · 被引用 28 次
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 被引用 27 次
- PivoTAL: Prior-Driven Supervision for Weakly-Supervised Temporal Action LocalizationMamshad Nayeem Rizve, Gaurav Mittal, Ye Yu, Matthew Hall 等CVPR 2023
- Boosting Weakly-Supervised Temporal Action Localization with Text InformationGuozhang Li, De Cheng, Xinpeng Ding, Nannan Wang 等CVPR 2023
- ASM-Loc: Action-aware Segment Modeling for Weakly-Supervised Temporal Action LocalizationBo He, Xitong Yang, Le Kang, Zhiyu Cheng 等CVPR 2022 · 被引用 104 次
