MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model
Shiyu Qin, XINJIE ZHANG, Zhening Liu, Jinpeng Wang, Bin Chen, Jiawei Li, Yifan Ren, Shu-Tao Xia, Jun Zhang
Abstract
Stereo image compression (SIC) has become increasingly vital with its applications surging in fields such as 3D reconstruction and autonomous navigation. Previous methods leverage cross-attention to model inter-view redundancy and employ autoregressive entropy models to predict probability distributions, achieving impressive rate-distortion performance. However, they suffer from slow coding speed due to the quadratic complexity of cross-attention mechanisms and the spatial autoregressive iterations of the entropy models. To address these limitations, we propose MambaSIC, which introduces two key innovations. First, we propose a Mamba-based stereo visual state space block (stereo VSSB) that leverages its linear complexity and long-range modeling capabilities to more rapidly and efficiently capture redundancy information between the two views. Second, to accelerate the compression process and enhance the accuracy of probability distribution estimation, we introduce a bi-directional multi-reference entropy model that utilizes a checkerboard partitioning strategy and the stereo VSSB to get rich inter-view priors. Experimental results demonstrate that our MambaSIC outperforms the state-of-the-art methods in both rate-distortion performance and coding efficiency. Moreover, it achieves the smallest inter-view PSNR discrepancy, resulting in more balanced reconstruction quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36c04c76-c07b-4a77-b18f-6f9b6b8e6f05Builds on19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- Transformer-based Transform CodingYinhao Zhu, Yang Yang, Taco CohenICLR 2022 · 218 citations
Related papers
- Cassic: Towards Content-Adaptive State-Space Models for Learned Image CompressionShiyu Qin, Jinpeng Wang, Yimin Zhou, Bin Chen et al.ICCV 2025 · 2 citations
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li et al.CVPR 2026
- CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image CompressionXinjie Zhang, Shenyuan Gao, Zhening Liu, Jiawei Shao et al.AAAI 2025 · 5 citations
- Deep Stereo Image Compression via Bi-directional CodingJianjun Lei, Xiangrui Liu, Bo Peng, Dengchao Jin et al.CVPR 2022 · 18 citations
- Disparity-based Stereo Image Compression with Aligned Cross-View PriorsYongqi Zhai, Luyang Tang, Yi Ma, Rui Peng et al.ACM MM 2022 · 10 citations
