EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding
Shuhan Tan, Tushar Nagarajan, Kristen Grauman
摘要
Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillation-based approach that learns to reconstruct heavy egocentric video clip features by combining the semantics from a sparse set of video frames with the head motion from lightweight IMU readings. We further devise a novel self-supervised training strategy for IMU feature learning. Our method leads to significant improvements in efficiency, requiring 200x fewer GFLOPs than equivalent video models. We demonstrate its effectiveness on the Ego4D and EPICKitchens datasets, where our method outperforms state-of-the-art efficient video understanding methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- UniMTS: Unified Pre-training for Motion Time SeriesXiyuan Zhang, Diyan Teng, Ranak Roy Chowdhury, Shuheng Li 等NeurIPS 2024 · 被引用 49 次
- Multimodal Distillation for Egocentric Action RecognitionGorjan Radevski, Dusan Grujicic, Matthew B. Blaschko, Marie-Francine Moens 等ICCV 2023 · 被引用 40 次
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu 等NeurIPS 2024 · 被引用 31 次
- EgoControl: Controllable Egocentric Video Generation via 3D Full-Body PosesEnrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 等CVPR 2026 · 被引用 7 次
- SkillSight: Efficient First-Person Skill Assessment with GazeChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 被引用 4 次
它引用的顶会 Paper22
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action RecognitionEvangelos Kazakos, Arsha Nagrani, Andrew Zisserman, Dima DamenICCV 2019 · 被引用 395 次
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra 等NeurIPS 2021 · 被引用 382 次
相关 Paper
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew 等UbiComp 2026 · 被引用 1 次
- Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric VideoYuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li 等CVPR 2026 · 被引用 1 次
- Ego-Exo: Transferring Visual Representations From Third-Person to First-Person VideosYanghao Li, Tushar Nagarajan, Bo Xiong, Kristen GraumanCVPR 2021
- DistInit: Learning Video Representations Without a Single Labeled VideoRohit Girdhar, Du Tran, Lorenzo Torresani, Deva RamananICCV 2019 · 被引用 59 次
- Self-Supervised Monocular 4D Scene Reconstruction for Egocentric VideosChengbo Yuan, Geng Chen, Li Yi, Yang GaoICCV 2025 · 被引用 2 次
