Neural Video Compression with Spatio-Temporal Cross-Covariance Transformers
Zhenghao Chen, Lucas Relic, Roberto Azevedo, Yang Zhang, Markus Gross, Dong Xu, Luping Zhou, Christopher Schroers
Abstract
Although existing neural video compression (NVC) methods have achieved significant success, most of them focus on improving either temporal or spatial information separately. They generally use simple operations such as concatenation or subtraction to utilize this information, while such operations only partially exploit spatio-temporal redundancies. This work aims to effectively and jointly leverage robust temporal and spatial information by proposing a new 3D-based transformer module: Spatio-Temporal Cross-Covariance Transformer (ST-XCT). The ST-XCT module combines two individual extracted features into a joint spatio-temporal feature, followed by 3D convolutional operations and a novel spatio-temporal-aware cross-covariance attention mechanism. Unlike conventional transformers, the cross-covariance attention mechanism is applied across the feature channels without breaking down the spatio-temporal features into local tokens. Such design allows for modeling global cross-channel correlations of the spatio-temporal context while lowering the computational requirement. Based on ST-XCT, we introduce a novel transformer-based end-to-end optimized NVC framework. ST-XCT-based modules are integrated into various key coding components of NVC, such as feature extraction, frame reconstruction, and entropy modeling, demonstrating its generalizability. Extensive experiments show that our ST-XCT-based NVC proposal achieves state-of-the-art compression performances on various standard video benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video CompressionZhenghao Chen, Luping Zhou, Zhihao Hu, Dong XuACM MM 2024 · 14 citations
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- 3D Gaussian Splatting Data Compression with Mixture of PriorsLei Liu, Zhenghao Chen, Dong XuACM MM 2025 · 4 citations
- EHVC: Efficient Hierarchical Reference and Quality Structure for Neural Video CodingJunqi Liao, Yaojun Wu, Chaoyi Lin, Zhipin Deng et al.ACM MM 2025 · 2 citations
- Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion ModelsYiming Wu, Zhenghao Chen, Huan Wang, Dong XuACM MM 2025 · 2 citations
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski et al.NeurIPS 2021 · 692 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Recurrent Video Restoration Transformer with Guided Deformable AttentionJingyun Liang, Yuchen Fan, Xiaoyu Xiang, Rakesh Ranjan et al.NeurIPS 2022 · 318 citations
Related papers
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai et al.ICCV 2021 · 224 citations
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
- UniFormer: Unified Transformer for Efficient Spatial-Temporal Representation LearningKunchang Li, Yali Wang, Peng Gao, Guanglu Song et al.ICLR 2022
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi et al.CVPR 2021
