Towards Practical Real-Time Neural Video Compression
Zhaoyang Jia, Bin Li, Jiahao Li, Wenxuan Xie, Linfeng Qi, Houqiang Li, Yan Lu
Abstract
We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) noncomputational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC . * This work was done when Zhaoyang Jia and Linfeng Qi were full-time interns at Microsoft Research Asia. † This paper is the outcome of an open-source project started from Dec. 2023.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb8d12bf-070d-4a74-88cb-f0b4c3bfc685Cited by top-tier papers21
- Generative Neural Video Compression via Video Diffusion PriorQi Mao, Hao Cheng, Tinghan Yang, Libiao Jin et al.CVPR 2026 · 18 citations
- Single-step Diffusion-based Video Coding with Semantic-Temporal GuidanceNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.CVPR 2026 · 12 citations
- Learned Image Compression with Hierarchical Progressive Context ModelingYuqi Li, Haotian Zhang, Li Li, Dong LiuICCV 2025 · 8 citations
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- Neural B-frame Video Compression with Bi-directional Reference HarmonizationYuxi Liu, Dengchao Jin, Shuai Huo, Jiawen Gu et al.NeurIPS 2025 · 6 citations
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair et al.ICCV 2021 · 137 citations
Related papers
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 20 citations
- Real-Time Neural Video Compression with Unified Intra and Inter CodingHui Xiang, Yifan Bian, Li Li, Jingran Wu et al.CVPR 2026 · 5 citations
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li et al.CVPR 2025
- Neural Video Compression with Feature ModulationJiahao Li, Bin Li, Yan LuCVPR 2024
- Neural Video Compression with Diverse ContextsJiahao Li, Bin Li, Yan LuCVPR 2023
