Augmented Deep Contexts for Spatially Embedded Video Coding
Yifan Bian, Chuanbo Tang, Li Li, Dong Liu
Abstract
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-ofthe-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https: //github.com/EsakaK/SEVC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Neural B-frame Video Compression with Bi-directional Reference HarmonizationYuxi Liu, Dengchao Jin, Shuai Huo, Jiawen Gu et al.NeurIPS 2025 · 6 citations
- Real-Time Neural Video Compression with Unified Intra and Inter CodingHui Xiang, Yifan Bian, Li Li, Jingran Wu et al.CVPR 2026 · 5 citations
- EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng et al.ACM MM 2025 · 2 citations
- Neural Video Compression with Reference HierarchyChuanbo Tang, Zhuoyuan Li, Li Li, Dong Liu et al.AAAI 2026
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren et al.NeurIPS 2021 · 430 citations
Related papers
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li et al.CVPR 2025
- Content-Adaptive Hierarchical Hyperprior for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng et al.CVPR 2026
- High Resolution Neural Video Coding with Bi-directional Confidence-Guided Reference Information ModelingFeng Ye, Kai Zhang, Li Zhang, Chuanmin JiaCVPR 2026
- Motion Information Propagation for Neural Video CompressionLinfeng Qi, Jiahao Li, Bin Li, Houqiang Li et al.CVPR 2023
