Augmented Deep Contexts for Spatially Embedded Video Coding
Yifan Bian, Chuanbo Tang, Li Li, Dong Liu
摘要
Most Neural Video Codecs (NVCs) only employ temporal references to generate temporal-only contexts and latent prior. These temporal-only NVCs fail to handle large motions or emerging objects due to limited contexts and misaligned latent prior. To relieve the limitations, we propose a Spatially Embedded Video Codec (SEVC), in which the low-resolution video is compressed for spatial references. Firstly, our SEVC leverages both spatial and temporal references to generate augmented motion vectors and hybrid spatial-temporal contexts. Secondly, to address the misalignment issue in latent prior and enrich the prior information, we introduce a spatial-guided latent prior augmented by multiple temporal latent representations. At last, we design a joint spatial-temporal optimization to learn quality-adaptive bit allocation for spatial references, further boosting rate-distortion performance. Experimental results show that our SEVC effectively alleviates the limitations in handling large motions or emerging objects, and also reduces 11.9% more bitrate than the previous state-ofthe-art NVC while providing an additional low-resolution bitstream. Our code and model are available at https: //github.com/EsakaK/SEVC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Neural B-frame Video Compression with Bi-directional Reference HarmonizationYuxi Liu, Dengchao Jin, Shuai Huo, Jiawen Gu 等NeurIPS 2025 · 被引用 6 次
- Real-Time Neural Video Compression with Unified Intra and Inter CodingHui Xiang, Yifan Bian, Li Li, Jingran Wu 等CVPR 2026 · 被引用 5 次
- EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng 等ACM MM 2025 · 被引用 2 次
- Neural Video Compression with Reference HierarchyChuanbo Tang, Zhuoyuan Li, Li Li, Dong Liu 等AAAI 2026
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- NeRV: Neural Representations for VideosHao Chen, Bo He, Hanyu Wang, Yixuan Ren 等NeurIPS 2021 · 被引用 430 次
相关 Paper
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 被引用 202 次
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li 等CVPR 2025
- Content-Adaptive Hierarchical Hyperprior for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng 等CVPR 2026
- High Resolution Neural Video Coding with Bi-directional Confidence-Guided Reference Information ModelingFeng Ye, Kai Zhang, Li Zhang, Chuanmin JiaCVPR 2026
- Motion Information Propagation for Neural Video CompressionLinfeng Qi, Jiahao Li, Bin Li, Houqiang Li 等CVPR 2023
