Neural Video Compression with Spatio-Temporal Cross-Covariance Transformers
Zhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang, Markus Gross, Dong Xu, Luping Zhou, Christopher Schroers
摘要
Although existing neural video compression (NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video CompressionZhenghao Chen, Luping Zhou, Zhihao Hu, Dong XuACM MM 2024 · 被引用 14 次
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li 等CVPR 2026 · 被引用 7 次
- 3D Gaussian Splatting Data Compression with Mixture of PriorsLei Liu, Zhenghao Chen, Dong XuACM MM 2025 · 被引用 4 次
- EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng 等ACM MM 2025 · 被引用 2 次
- Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion ModelsYiming Wu, Zhenghao Chen, Huan Wang, Dong XuACM MM 2025 · 被引用 2 次
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski 等NeurIPS 2021 · 被引用 692 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan 等NeurIPS 2022 · 被引用 318 次
相关 Paper
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai 等ICCV 2021 · 被引用 224 次
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song 等ICLR 2022
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi 等CVPR 2021
