Rethinking Resolution in the Context of Efficient Video Recognition
Chuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo, Zehuan Yuan, Xiaojuan Qi
摘要
In this paper, we empirically study how to make the most of low-resolution frames for efficient video recognition. Existing methods mainly focus on developing compact networks or alleviating temporal redundancy of video inputs to increase efficiency, whereas compressing frame resolution has rarely been considered a promising solution. A major concern is the poor recognition accuracy on lowresolution frames. We thus start by analyzing the underlying causes of performance degradation on low-resolution frames. Our key finding is that the major cause of degradation is not information loss in the down-sampling process, but rather the mismatch between network architecture and input scale. Motivated by the success of knowledge distillation (KD), we propose to bridge the gap between network and input size via cross-resolution KD (ResKD). Our work shows that ResKD is a simple but effective method to boost recognition accuracy on low-resolution frames. Without bells and whistles, ResKD considerably surpasses all competitive methods in terms of efficiency and accuracy on four large-scale benchmark datasets, i.e., ActivityNet, FCVID, Mini-Kinetics, Something-Something V2. In addition, we extensively demonstrate its effectiveness over state-of-the-art architectures, i.e., 3D-CNNs and Video Transformers, and scalability towards super low-resolution frames. The results suggest ResKD can serve as a general inference acceleration method for state-of-the-art video recognition. Our code will be available at https://github.com/CVMI-Lab/ResKD . * This work was performed when Chuofan Ma worked as an intern at ByteDance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OmniViD: A Generative Framework for Universal Video UnderstandingJunke Wang, Dongdong Chen, Chong Luo, Bo He 等CVPR 2024 · 被引用 18 次
- A Simple Recipe for Contrastively Pre-Training Video-First Encoders Beyond 16 FramesPinelopi Papalampidi, Skanda Koppula, Shreya Pathak, Justin Chiu 等CVPR 2024 · 被引用 15 次
- Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosYubin Hu, Yuze He, Yanghao Li, Jisheng Li 等CVPR 2023
- ResFormer: Scaling ViTs with Multi-Resolution TrainingRui Tian, Zuxuan Wu, Qi Dai, Han Hu 等CVPR 2023
它引用的顶会 Paper22
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
相关 Paper
- Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video ModelsYing Peng, Hongsen Ye, Changxin Huang, Xiping Hu 等AAAI 2026
- Ultrafast Video Attention Prediction with Coupled Knowledge DistillationKui Fu, Peipei Shi, Yafei Song, Shiming Ge 等AAAI 2020 · 被引用 11 次
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao 等AAAI 2024 · 被引用 9 次
- Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation LearningRui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen 等CVPR 2023
- ResidualViT for Efficient Temporally Dense Video EncodingMattia Soldan, Fabian Caba Heilbron, Bernard Ghanem, Josef Sivic 等ICCV 2025
