Ultrafast Video Attention Prediction with Coupled Knowledge Distillation
Kui Fu, Peipei Shi, Yafei Song, Shiming Ge, Xiangju Lu, Jia Li
Abstract
Large convolutional neural network models have recently demonstrated impressive performance on video attention prediction. Conventionally, these models are with intensive computation and large memory. To address these issues, we design an extremely light-weight network with ultrafast speed, named UVA-Net. The network is constructed based on depth-wise convolutions and takes low-resolution images as input. However, this straight-forward acceleration method will decrease performance dramatically. To this end, we propose a coupled knowledge distillation strategy to augment and train the network effectively. With this strategy, the model can further automatically discover and emphasize implicit useful cues contained in the data. Both spatial and temporal knowledge learned by the high-resolution complex teacher networks also can be distilled and transferred into the proposed low-resolution light-weight spatiotemporal network. Experimental results show that the performance of our model is comparable to 11 state-of-the-art models in video attention prediction, while it costs only 0.68 MB memory footprint, runs about 10,106 FPS on GPU and 404 FPS on CPU, which is 206 times faster than previous models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Temporally Distributed Networks for Fast Video Semantic SegmentationPing Hu, Fabian Caba, Oliver Wang, Zhe Lin et al.CVPR 2020
- Space-Time Distillation for Video Super-ResolutionZeyu Xiao, Xueyang Fu, Jie Huang, Zhen Cheng et al.CVPR 2021
- Lightweight Spatio-Temporal Modeling via Temporally Shifted Distillation for Real-Time Accident AnticipationPatrik Patera, Yie-Tarng Chen, Wen-Hsien FangICLR 2026
- Rethinking Resolution in the Context of Efficient Video RecognitionChuofan Ma, Qiushan Guo, Yi Jiang, Ping Luo et al.NeurIPS 2022 · 17 citations
- Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal ForecastingYuqi Li, Chuanguang Yang, Hansheng Zeng, Zeyu Dong et al.ICCV 2025 · 23 citations
