Two-Stream Action Recognition-Oriented Video Super-Resolution
Haochen Zhang, Dong Liu, Zhiwei Xiong
摘要
We study the video super-resolution (SR) problem for facilitating video analytics tasks, e.g. action recognition, instead of for visual quality. The popular action recognition methods based on convolutional networks, exemplified by two-stream networks, are not directly applicable on video of low spatial resolution. This can be remedied by performing video SR prior to recognition, which motivates us to improve the SR procedure for recognition accuracy. Tailored for two-stream action recognition networks, we propose two video SR methods for the spatial and temporal streams respectively. On the one hand, we observe that regions with action are more important to recognition, and we propose an optical-flow guided weighted mean-squared-error loss for our spatial-oriented SR (SoSR) network to emphasize the reconstruction of moving objects. On the other hand, we observe that existing video SR methods incur temporal discontinuity between frames, which also worsens the recognition accuracy, and we propose a siamese network for our temporal-oriented SR (ToSR) training that emphasizes the temporal continuity between consecutive frames. We perform experiments using two state-of-the-art action recognition networks and two well-known datasets--UCF101 and HMDB51. Results demonstrate the effectiveness of our proposed SoSR and ToSR in improving recognition accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Reconstructed Convolution Module Based Look-Up Tables for Efficient Image Super-ResolutionGuandu Liu, Yukang Ding, Mading Li, Ming Sun 等ICCV 2023 · 被引用 25 次
- Turning Frequency to Resolution: Video Super-Resolution via Event CamerasYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song 等CVPR 2021
- Space-Time Distillation for Video Super-ResolutionZeyu Xiao, Xueyang Fu, Jie Huang, Zhen Cheng 等CVPR 2021
- Learning to Have an Ear for Face Super-ResolutionGivi Meishvili, Simon Jenni, Paolo FavaroCVPR 2020
- Structured Sparsity Learning for Efficient Video Super-ResolutionBin Xia, Jingwen He, Yulun Zhang, Yitong Wang 等CVPR 2023
相关 Paper
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 被引用 41 次
- A Slow-I-Fast-P Architecture for Compressed Video Action RecognitionJiapeng Li, Ping Wei, Yongchi Zhang, Nanning ZhengACM MM 2020 · 被引用 49 次
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij 等CVPR 2021
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action RecognitionIshan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak ShahCVPR 2023
