Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding
Peijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex C. Kot
摘要
This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently compromise performance due to inadequate supervision. Therefore, to tackle this challenge, we for the first time pay attention to exploiting complementary information extracted from multi-modal videos (e.g., RGB frames, optical flows), where richer supervision is naturally introduced in the weaklysupervised context. Our motivation is that by integrating different modalities of the videos, the model is learned from synergic supervision and thereby can attain superior generalization capability. However, addressing multiple modalities would also inevitably introduce additional computational overhead, and might become inapplicable if a particular modality is inaccessible. To solve this issue, we adopt a novel route: building a multi-modal distillation algorithm to capitalize on the multi-modal knowledge as supervision for model training, while still being able to work with only the single modal input during inference. As such, we can utilize the benefits brought by the supplementary nature of multiple modalities, without undermining the applicability in practical scenarios. Specifically, we first propose a cross-modal mutual learning framework and train a sophisticated teacher model to learn collaboratively from the multi-modal videos. Then we identify two sorts of knowledge from the teacher model, i.e., temporal boundaries and semantic activation map. And we devise a local-global distillation algorithm to transfer this knowledge to a student model of single-modal input at both local and global levels. Extensive experiments on large-scale datasets demonstrate that our method achieves state-of-the-art performance with/without multi-modal inputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets ConsistencyPeijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng 等AAAI 2024 · 被引用 11 次
- Vid-Group: Temporal Video Grounding Pretraining from Unlabeled Videos in the WildPeijun Bao, Chenqi Kong, Siyuan Yang, Zihao Shao 等ICCV 2025 · 被引用 3 次
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosPeijun Bao, Anwei Luo, Gang Pan, Alex C. Kot 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper15
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Temporally Grounding Language Queries in Videos by Contextual Boundary-Aware PredictionJingwen Wang, Lin Ma, Wenhao JiangAAAI 2020 · 被引用 206 次
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 等CVPR 2022 · 被引用 108 次
- Regularized Two-Branch Proposal Networks for Weakly-Supervised Moment Retrieval in VideosZhu Zhang, Zhijie Lin, Zhou Zhao, Jieming Zhu 等ACM MM 2020 · 被引用 86 次
相关 Paper
- End-to-end Multi-modal Video Temporal GroundingYi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan YangNeurIPS 2021 · 被引用 68 次
- Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionPilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee 等CVPR 2023
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 被引用 50 次
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 被引用 27 次
- Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph GroundingChaolei Tan, Jianhuang Lai, Wei-Shi Zheng, Jian-Fang HuCVPR 2024 · 被引用 5 次
