Disparity-based Stereo Image Compression with Aligned Cross-View Priors
Yongqi Zhai, Luyang Tang, Yi Ma, Rui Peng, Ronggang Wang
Abstract
With the wide application of stereo images in various fields, the research on stereo image compression (SIC) attracts extensive attention from academia and industry. The core of SIC is to fully explore the mutual information between the left and right images and reduce redundancy between views as much as possible. In this paper, we propose DispSIC, an end-to-end trainable deep neural network, in which we jointly train a stereo matching model to assist in the image compression task. Based on the stereo matching results (i.e. disparity), the right image can be easily warped to the left view, and only the residuals between the left and right views are encoded for the left image. A three-branch auto-encoder architecture is adopted in DispSIC, which encodes the right image, the disparity map and the residuals respectively. During training, the whole network can learn how to adaptively allocate bitrates to these three parts, achieving better rate-distortion performance at the cost of a lower disparity map bitrates. Moreover, we propose a conditional entropy model with aligned cross-view priors for SIC, which takes the warped latents of the right image as priors to improve the accuracy of the probability estimation for the left image. Experimental results demonstrate that our proposed method achieves superior performance compared to other existing SIC methods on the KITTI and InStereo2K datasets both quantitatively and qualitatively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07583746-37ac-4c71-8de5-169ee0c5fcafCited by top-tier papers5
- CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image CompressionXinjie Zhang, Shenyuan Gao, Zhening Liu, Jiawei Shao et al.AAAI 2025 · 5 citations
- MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy ModelShiyu Qin, XINJIE ZHANG, Zhening Liu, Jinpeng Wang et al.CVPR 2026
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li et al.CVPR 2026
- 3D-LMVIC: Learning-based Multi-View Image Compression with 3D Gaussian Geometric PriorsYujun Huang, Bin Chen, Niu Lian, Xin Wang et al.ICML 2025
- SDiD:Shared diffusion prior for efficient distributed stereo image compressionYichong Xia, Yimin Zhou, Zongyu Li, Shiyu Qin et al.ICML 2026
Builds on14
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 675 citations
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Variable Rate Deep Image Compression With a Conditional AutoencoderYoojin Choi, Mostafa El-Khamy, Jungwon LeeICCV 2019 · 265 citations
- Transformer-based Transform CodingYinhao Zhu, Yang Yang, Taco CohenICLR 2022 · 218 citations
- Enhanced Invertible Encoding for Learned Image CompressionYueqi Xie, Ka Leong Cheng, Qifeng ChenACM MM 2021 · 195 citations
Related papers
- Deep Homography for Efficient Stereo Image CompressionXin Deng, Wenzhe Yang, Ren Yang, Mai Xu et al.CVPR 2021
- Deep Stereo Image Compression via Bi-directional CodingJianjun Lei, Xiangrui Liu, Bo Peng, Dengchao Jin et al.CVPR 2022 · 18 citations
- SASIC: Stereo Image Compression with Latent Shifts and Stereo AttentionMatthias Wödlinger, Jan Kotera, Jan Xu, Robert SablatnigCVPR 2022 · 24 citations
- DSIC: Deep Stereo Image CompressionJerry Liu, Shenlong Wang, Raquel UrtasunICCV 2019 · 50 citations
- LSVC: A Learning-based Stereo Video Compression FrameworkZhenghao Chen, Guo Lu, Zhihao Hu, Shan Liu et al.CVPR 2022 · 39 citations
