Multimodal Distillation for Egocentric Action Recognition
Gorjan Radevski, Dusan Grujicic, Matthew B. Blaschko, Marie-Francine Moens, Tinne Tuytelaars
摘要
The focal point of egocentric video understanding is modelling hand-object interactions. Standard models, e.g. CNNs or Vision Transformers, which receive RGB frames as input perform well, however, their performance improves further by employing additional input modalities (e.g. object detections, optical flow, audio, etc.) which provide cues complementary to the RGB modality. The added complexity of the modality-specific modules, on the other hand, makes these models impractical for deployment. The goal of this work is to retain the performance of such a multi-modal approach, while using only the RGB frames as input at inference time. We demonstrate that for egocentric action recognition on the Epic-Kitchens and the Something-Something datasets, students which are taught by multi-modal teachers tend to be more accurate and better calibrated than architecturally equivalent models trained on ground truth labels in a unimodal or multimodal fashion. We further adopt a principled multimodal knowledge distillation framework, allowing us to deal with issues which occur when applying multimodal knowledge distillation in a naïve manner. Lastly, we demonstrate the achieved reduction in computational complexity, and show that our approach maintains higher performance with the reduction of the number of input views. We release our code at: https://github.com/gorjanradevski/multimodal-distillation
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song 等ICML 2024 · 被引用 23 次
- Balancing Multimodal Training Through Game-Theoretic RegularizationKonstantinos Kontras, Thomas Strypsteen, Christos Chatzichristos, Paul Pu Liang 等NeurIPS 2025 · 被引用 17 次
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu 等AAAI 2025 · 被引用 8 次
- Compositional Steering of Large Language Models with Steering TokensGorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence 等ACL 2026 · 被引用 4 次
- Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue ConsistencyZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
相关 Paper
- Multimodal Global Relation Knowledge Distillation for Egocentric Action AnticipationYi Huang, Xiaoshan Yang, Changsheng XuACM MM 2021 · 被引用 11 次
- Interact before Align: Leveraging Cross-Modal Knowledge for Domain Adaptive Action RecognitionLijin Yang, Yifei Huang, Yusuke Sugano, Yoichi SatoCVPR 2022 · 被引用 35 次
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 被引用 50 次
- Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionPilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee 等CVPR 2023
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
