Scale-Space Flow for End-to-End Optimized Video Compression
Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Ballé, Sung Jin Hwang, George Toderici
Abstract
Despite considerable progress on end-to-end optimized deep networks for image compression, video coding remains a challenging task. Recently proposed methods for learned video compression use optical flow and bilinear warping for motion compensation and show competitive rate-distortion performance relative to hand-engineered codecs like H.264 and HEVC. However, these learningbased methods rely on complex architectures and training schemes including the use of pre-trained optical flow networks, disjoint training of sub-networks, adaptive rate control, and buffering intermediate reconstructions to disk during training. In this paper, we show that a generalized warping operator that better handles common failure cases, e.g. disocclusions and fast motion, can provide competitive compression results with a greatly simplified model and training procedure. Specifically, we propose scale-space flow, an intuitive generalization of 2D optical flow that adds a scale parameter to allow the network to better model uncertainty. Our experiments show that a low-latency video compression model (no B-frames) using scale-space flow for motion compensation can outperform analogous stateof-the art learned video compression models while being trained using a much simpler procedure and without any pre-trained optical flow networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a49f3e3-0c9d-4ef3-9871-9f64454e720bCited by top-tier papers82
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
- Transformer-based Transform CodingYinhao Zhu, Yang Yang, Taco CohenICLR 2022 · 218 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
Builds on5
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Variable Rate Deep Image Compression With a Conditional AutoencoderYoojin Choi, Mostafa El-Khamy, Jungwon LeeICCV 2019 · 265 citations
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson et al.ICCV 2019 · 258 citations
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
Related papers
- Offline and Online Optical Flow Enhancement for Deep Video CompressionChuanbo Tang, Xihua Sheng, Zhuoyuan Li, Haotian Zhang et al.AAAI 2024 · 35 citations
- Learning Video Stabilization Using Optical FlowJiyang Yu, Ravi RamamoorthiCVPR 2020
- M-LVC: Multiple Frames Prediction for Learned Video CompressionJianping Lin, Dong Liu, Houqiang Li, Feng WuCVPR 2020
- Hierarchical B-Frame Video Coding Using Two-Layer CANF Without Motion CodingDavid Alexandre, Hsueh-Ming Hang, Wen-Hsiao PengCVPR 2023
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair et al.ICCV 2021 · 137 citations
