Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action Detection
Rui Dai, Srijan Das, François Brémond
摘要
In video understanding, most cross-modal knowledge distillation (KD) methods are tailored for classification tasks, focusing on the discriminative representation of the trimmed videos. However, action detection requires not only categorizing actions, but also localizing them in untrimmed videos. Therefore, transferring knowledge pertaining to temporal relations is critical for this task which is missing in the previous cross-modal KD frameworks. To this end, we aim at learning an augmented RGB representation for action detection, taking advantage of additional modalities at training time through KD. We propose a KD frame-work consisting of two levels of distillation. On one hand, atomic-level distillation encourages the RGB student to learn the sub-representation of the actions from the teacher in a contrastive manner. On the other hand, sequence-level distillation encourages the student to learn the temporal knowledge from the teacher, which consists of transferring the Global Contextual Relations and the Action Boundary Saliency. The result is an Augmented-RGB stream that can achieve competitive performance as the two-stream network while using only RGB at inference time. Extensive experimental analysis shows that our proposed distillation frame-work is generic and outperforms other popular cross-modal distillation methods in action detection task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionRui Dai, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo 等CVPR 2022 · 被引用 93 次
- XKD: Cross-Modal Knowledge Distillation with Domain Alignment for Video Representation LearningPritam Sarkar, Ali EtemadAAAI 2024 · 被引用 45 次
- LAC - Latent Action Composition for Skeleton-based Action SegmentationDi Yang, Yaohui Wang, Antitza Dantcheva, Quan Kong 等ICCV 2023 · 被引用 22 次
- CroMo: Cross-Modal Learning for Monocular Depth EstimationYannick Verdié, Jifei Song, Barnabé Mas, Benjamin Busam 等CVPR 2022 · 被引用 17 次
- DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action LocalizationXiaojun Tang, Junsong Fan, Chuanchen Luo, Zhaoxiang Zhang 等ICCV 2023 · 被引用 16 次
它引用的顶会 Paper6
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- Learning Salient Boundary Feature for Anchor-free Temporal Action LocalizationChuming Lin, Chengming Xu, Donghao Luo, Yabiao Wang 等CVPR 2021
- G-TAD: Sub-Graph Localization for Temporal Action DetectionMengmeng Xu, Chen Zhao, David S. Rojas, Ali K. Thabet 等CVPR 2020
相关 Paper
- Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionPilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee 等CVPR 2023
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao 等AAAI 2024 · 被引用 9 次
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 被引用 11 次
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingPeijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng 等AAAI 2024 · 被引用 20 次
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 被引用 2 次
