Two-Stream Action Recognition-Oriented Video Super-Resolution
Haochen Zhang, Dong Liu, Zhiwei Xiong
Abstract
We study the video super-resolution (SR) problem for facilitating video analytics tasks, e.g. action recognition, instead of for visual quality. The popular action recognition methods based on convolutional networks, exemplified by two-stream networks, are not directly applicable on video of low spatial resolution. This can be remedied by performing video SR prior to recognition, which motivates us to improve the SR procedure for recognition accuracy. Tailored for two-stream action recognition networks, we propose two video SR methods for the spatial and temporal streams respectively. On the one hand, we observe that regions with action are more important to recognition, and we propose an optical-flow guided weighted mean-squared-error loss for our spatial-oriented SR (SoSR) network to emphasize the reconstruction of moving objects. On the other hand, we observe that existing video SR methods incur temporal discontinuity between frames, which also worsens the recognition accuracy, and we propose a siamese network for our temporal-oriented SR (ToSR) training that emphasizes the temporal continuity between consecutive frames. We perform experiments using two state-of-the-art action recognition networks and two well-known datasets--UCF101 and HMDB51. Results demonstrate the effectiveness of our proposed SoSR and ToSR in improving recognition accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e23bd494-a95b-4fcc-a14c-fda6e9e1d33fCited by top-tier papers7
- Reconstructed Convolution Module Based Look-Up Tables for Efficient Image Super-ResolutionGuandu Liu, Yukang Ding, Mading Li, Ming Sun et al.ICCV 2023 · 25 citations
- Turning Frequency to Resolution: Video Super-Resolution via Event CamerasYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.CVPR 2021
- Space-Time Distillation for Video Super-ResolutionZeyu Xiao, Xueyang Fu, Jie Huang, Zhen Cheng et al.CVPR 2021
- Learning to Have an Ear for Face Super-ResolutionGivi Meishvili, Simon Jenni, Paolo FavaroCVPR 2020
- Structured Sparsity Learning for Efficient Video Super-ResolutionBin Xia, Jingwen He, Yulun Zhang, Yitong Wang et al.CVPR 2023
Related papers
- Hierarchical Self-Attention Network for Action Localization in VideosRizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien FangICCV 2019 · 41 citations
- A Slow-I-Fast-P Architecture for Compressed Video Action RecognitionJiapeng Li, Ping Wei, Yongchi Zhang, Nanning ZhengACM MM 2020 · 49 citations
- No Frame Left Behind: Full Video Action RecognitionXin Liu, Silvia L. Pintea, Fatemeh Karimi Nejadasl, Olaf Booij et al.CVPR 2021
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action RecognitionIshan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak ShahCVPR 2023
