Decomposed Cross-Modal Distillation for RGB-based Temporal Action Detection
Pilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee, Hyeran Byun
摘要
Temporal action detection aims to predict the time intervals and the classes of action instances in the video. Despite the promising performance, existing two-stream models exhibit slow inference speed due to their reliance on computationally expensive optical flow. In this paper, we introduce a decomposed cross-modal distillation framework to build a strong RGB-based detector by transferring knowledge of the motion modality. Specifically, instead of direct distillation, we propose to separately learn RGB and motion representations, which are in turn combined to perform action localization. The dual-branch design and the asymmetric training objectives enable effective motion knowledge transfer while preserving RGB information intact. In addition, we introduce a local attentive fusion to better exploit the multimodal complementarity. It is designed to preserve the local discriminability of the features that is important for action localization. Extensive experiments on the benchmarks verify the effectiveness of the proposed method in enhancing RGB-based action detectors. Notably, our framework is agnostic to backbones and detection heads, bringing consistent gains across different model combinations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Breaking Modality Gap in RGBT Tracking: Coupled Knowledge DistillationAndong Lu, Jiacong Zhao, Chenglong Li, Yun Xiao 等ACM MM 2024 · 被引用 15 次
- Realigning Confidence with Temporal Saliency Information for Point-Level Weakly-Supervised Temporal Action LocalizationZiying Xia, Jian Cheng, Siyu Liu, Yongxiang Hu 等CVPR 2024 · 被引用 10 次
- FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal DataYiting Li, Fayao Liu, Jingyi Liao, Sichao Tian 等ICCV 2025 · 被引用 5 次
- Language Model Guided Interpretable Video Action ReasoningNing Wang, Guangming Zhu, HS Li, Liang Zhang 等CVPR 2024 · 被引用 3 次
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper37
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding 等ICCV 2019 · 被引用 709 次
相关 Paper
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 被引用 50 次
- Multimodal Decomposed Distillation with Instance Alignment and Uncertainty Compensation for Thermal Object DetectionYanfeng Liu, Lefei ZhangACM MM 2025 · 被引用 2 次
- Efficient RGB-T Tracking via Cross-Modality DistillationTianlu Zhang, Hongyuan Guo, Qiang Jiao, Qiang Zhang 等CVPR 2023
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingPeijun Bao, Yong Xia, Wenhan Yang, Boon Poh Ng 等AAAI 2024 · 被引用 20 次
- Weakly-Supervised Temporal Action Localization via Cross-Stream Collaborative LearningYuan Ji, Xu Jia, Huchuan Lu, Xiang RuanACM MM 2021 · 被引用 27 次
