MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model
Shiyu Qin, XINJIE ZHANG, Zhening Liu, Jinpeng Wang, Bin Chen, Jiawei Li, Yifan Ren, Shu-Tao Xia, Jun Zhang
摘要
Stereo image compression (SIC) has become increasingly vital with its applications surging in fields such as 3D reconstruction and autonomous navigation. Previous methods leverage cross-attention to model inter-view redundancy and employ autoregressive entropy models to predict probability distributions, achieving impressive rate-distortion performance. However, they suffer from slow coding speed due to the quadratic complexity of cross-attention mechanisms and the spatial autoregressive iterations of the entropy models. To address these limitations, we propose MambaSIC, which introduces two key innovations. First, we propose a Mamba-based stereo visual state space block (stereo VSSB) that leverages its linear complexity and long-range modeling capabilities to more rapidly and efficiently capture redundancy information between the two views. Second, to accelerate the compression process and enhance the accuracy of probability distribution estimation, we introduce a bi-directional multi-reference entropy model that utilizes a checkerboard partitioning strategy and the stereo VSSB to get rich inter-view priors. Experimental results demonstrate that our MambaSIC outperforms the state-of-the-art methods in both rate-distortion performance and coding efficiency. Moreover, it achieves the smallest inter-view PSNR discrepancy, resulting in more balanced reconstruction quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma 等CVPR 2022 · 被引用 363 次
- Transformer-based Transform CodingYinhao Zhu, Yang Yang, Taco CohenICLR 2022 · 被引用 218 次
相关 Paper
- Cassic: Towards Content-Adaptive State-Space Models for Learned Image CompressionShiyu Qin, Jinpeng Wang, Yimin Zhou, Bin Chen 等ICCV 2025 · 被引用 2 次
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li 等CVPR 2026
- CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image CompressionXinjie Zhang, Shenyuan Gao, Zhening Liu, Jiawei Shao 等AAAI 2025 · 被引用 5 次
- Deep Stereo Image Compression via Bi-directional CodingJianjun Lei, Xiangrui Liu, Bo Peng, Dengchao Jin 等CVPR 2022 · 被引用 18 次
- Disparity-based Stereo Image Compression with Aligned Cross-View PriorsYongqi Zhai, Luyang Tang, Yi Ma, Rui Peng 等ACM MM 2022 · 被引用 10 次
