Scene Matters: Model-based Deep Video Compression
Lv Tang, Xinfeng Zhang, Gai Zhang, Xiaoqi Ma
摘要
Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal prediction theory to enhance compression performance by designing high efficient intra and inter prediction strategies and compressing video frames one by one. In this paper, we propose a novel model-based video compression (MVC) framework that regards scenes as the fundamental units for video sequences. Our proposed MVC directly models the intensity variation of the entire video sequence in one scene, seeking non-redundant representations instead of reducing redundancy through spatio-temporal predictions. To achieve this, we employ implicit neural representation as our basic modeling architecture. To improve the efficiency of video modeling, we first propose context-related spatial positional embedding and frequency domain supervision in spatial context enhancement. For temporal correlation capturing, we design the scene flow constrain mechanism and temporal contrastive loss. Extensive experimental results demonstrate that our method achieves up to a 20% bitrate reduction compared to the latest video coding standard H.266 and is more efficient in decoding than existing video coding strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Boosting Neural Representations for Videos with a Conditional DecoderXinjie Zhang, Ren Yang, Dailan He, Xingtong Ge 等CVPR 2024 · 被引用 20 次
- Context Guided Transformer Entropy Modeling for Video CompressionJunlong Tong, Wei Zhang, Yaohui Jin, Xiaoyu ShenICCV 2025
- An Exploration with Entropy Constrained 3D Gaussians for 2D Video CompressionXiang Liu, Bin Chen, Zimo Liu, Yaowei Wang 等ICLR 2025
它引用的顶会 Paper21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren 等NeurIPS 2021 · 被引用 430 次
相关 Paper
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li 等CVPR 2025
- Neural Video Compression with Reference HierarchyChuanbo Tang, Zhuoyuan Li, Li Li, Dong Liu 等AAAI 2026
- FFNeRV: Flow-Guided Frame-Wise Neural Representations for VideosJoo Chan Lee, Daniel Rho, Jong Hwan Ko, Eunbyung ParkACM MM 2023 · 被引用 59 次
- HiNeRV: Video Compression with Hierarchical Encoding-based Neural RepresentationHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower 等NeurIPS 2023 · 被引用 132 次
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower 等NeurIPS 2024 · 被引用 44 次
