Towards Practical Real-Time Neural Video Compression
Zhaoyang Jia, Bin Li, Jiahao Li, Wenxuan Xie, Linfeng Qi, Houqiang Li, Yan Lu
摘要
We introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) noncomputational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC . * This work was done when Zhaoyang Jia and Linfeng Qi were full-time interns at Microsoft Research Asia. † This paper is the outcome of an open-source project started from Dec. 2023.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Generative Neural Video Compression via Video Diffusion PriorQi Mao, Hao Cheng, Tinghan Yang, Libiao Jin 等CVPR 2026 · 被引用 18 次
- Single-step Diffusion-based Video Coding with Semantic-Temporal GuidanceNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li 等CVPR 2026 · 被引用 12 次
- Learned Image Compression with Hierarchical Progressive Context ModelingYuqi Li, Haotian Zhang, Li Li, Dong LiuICCV 2025 · 被引用 8 次
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li 等CVPR 2026 · 被引用 7 次
- Neural B-frame Video Compression with Bi-directional Reference HarmonizationYuxi Liu, Dengchao Jin, Shuai Huo, Jiawen Gu 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 被引用 202 次
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles 等NeurIPS 2022 · 被引用 155 次
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair 等ICCV 2021 · 被引用 137 次
相关 Paper
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 被引用 20 次
- Real-Time Neural Video Compression with Unified Intra and Inter CodingHui Xiang, Yifan Bian, Li Li, Jingran Wu 等CVPR 2026 · 被引用 5 次
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li 等CVPR 2025
- Neural Video Compression with Feature ModulationJiahao Li, Bin Li, Yan LuCVPR 2024
- Neural Video Compression with Diverse ContextsJiahao Li, Bin Li, Yan LuCVPR 2023
