Towards High-Quality and Efficient Video Super-Resolution via Spatial-Temporal Data Overfitting
Gen Li, Jie Ji, Minghai Qin, Wei Niu, Bin Ren, Fatemeh Afghah, Linke Guo, Xiaolong Ma
Abstract
As deep convolutional neural networks (DNNs) are widely used in various fields of computer vision, leveraging the overfitting ability of the DNN to achieve video resolution upscaling has become a new trend in the modern video delivery system. By dividing videos into chunks and overfitting each chunk with a super-resolution model, the server encodes videos before transmitting them to the clients, thus achieving better video quality and transmission efficiency. However, a large number of chunks are expected to ensure good overfitting quality, which substantially increases the storage and consumes more bandwidth resources for data transmission. On the other hand, decreasing the number of chunks through training optimization techniques usually requires high model capacity, which significantly slows down execution speed. To reconcile such, we propose a novel method for high-quality and efficient video resolution upscaling tasks, which leverages the spatial-temporal information to accurately divide video into chunks, thus keeping the number of chunks as well as the model size to minimum. Additionally, we advance our method into a single overfitting model by a data-aware joint training technique, which further reduces the storage requirement with negligible quality drop. We deploy our models on an offthe-shelf mobile phone, and experimental results show that our method achieves real-time video super-resolution with high video quality. Compared with the state-of-the-art, our method achieves 28 fps streaming speed with 41.6 PSNR, which is 14× faster and 2.29 dB better in the live video resolution upscaling tasks. Code available in https:// github.com/coulsonlee/STDO-CVPR2023.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- 4KAgent: Agentic Any Image to 4K Super-ResolutionYushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang et al.NeurIPS 2025 · 51 citations
- Advancing Dynamic Sparse Training by Exploring Optimization OpportunitiesJie Ji, Gen Li, Lu Yin, Minghai Qin et al.ICML 2024 · 10 citations
- NeurRev: Train Better Sparse Neural Network Practically via Neuron RevitalizationGen Li, Lu Yin, Jie Ji, Wei Niu et al.ICLR 2024 · 10 citations
- Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolutionQi Tang, Yao Zhao, Meiqin Liu, Jian Jin et al.AAAI 2024 · 10 citations
- Low-Latency Space-Time Supersampling for Real-Time RenderingRuian He, Shili Zhou, Yuqi Sun, Ri Cheng et al.AAAI 2024 · 5 citations
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- FcaNet: Frequency Channel Attention NetworksZequn Qin, Pengyi Zhang, Fei Wu, Xi LiICCV 2021 · 1,049 citations
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson et al.ICCV 2019 · 258 citations
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
Related papers
- Overfitting the Data: Compact Neural Video Delivery via Content-aware Feature ModulationJiaming Liu, Ming Lu, Kaixin Chen, Xiaoqi Li et al.ICCV 2021 · 39 citations
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 68 citations
- E2SR: an end-to-end video CODEC assisted system for super resolution accelerationZhuoran Song, Zhongkai Yu, Naifeng Jing, Xiaoyao LiangDAC 2022 · 4 citations
- GameStreamSR: Enabling Neural-Augmented Game Streaming on Commodity Mobile PlatformsSandeepa Bhuyan, Ziyu Ying, Mahmut T. Kandemir, Mahanth Gowda et al.ISCA 2024 · 6 citations
- BiSR: Bidirectionally Optimized Super-Resolution for Mobile Video StreamingQian Yu, Qing Li, Rui He, Gareth Tyson et al.WWW 2023 · 11 citations
