CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
Xinjie Zhang, Shenyuan Gao, Zhening Liu, Jiawei Shao, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang, Jun Zhang
Abstract
Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, which leads to suboptimal rate-distortion results. In this paper, we propose a stereo image compression framework, named CAMSIC. CAMSIC independently transforms each image to latent representation and employs a powerful decoder-free Transformer entropy model to capture both spatial and disparity dependencies, by introducing a novel content-aware masked image modeling (MIM) technique. Our content-aware MIM facilitates efficient bidirectional interaction between prior information and estimated tokens, which naturally obviates the need for an extra Transformer decoder. Experiments show that our stereo image codec achieves state-of-the-art rate-distortion performance on two stereo image datasets Cityscapes and InStereo2K with fast encoding and decoding speed. * This work was partially performed when Xinjie Zhang was an Intern at SenseTime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee428ec5-3686-4f90-84a3-6304bcf69bdaCited by top-tier papers5
- MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy ModelShiyu Qin, XINJIE ZHANG, Zhening Liu, Jinpeng Wang et al.CVPR 2026
- Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image CompressionHaotian Zhang, Feiyue Long, Yixin Yu, Jian Xue et al.CVPR 2026
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li et al.CVPR 2026
- Distributed Image Compression with Multimodal Side Information at Extremely Low BitratesGuojun Xu, Mingyang Zhang, Jianwen Xiang, Cheng Tan et al.CVPR 2026
- SDiD:Shared diffusion prior for efficient distributed stereo image compressionYichong Xia, Yimin Zhou, Zongyu Li, Shiyu Qin et al.ICML 2026
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 690 citations
Related papers
- Deep Stereo Image Compression via Bi-directional CodingJianjun Lei, Xiangrui Liu, Bo Peng, Dengchao Jin et al.CVPR 2022 · 18 citations
- Disparity-based Stereo Image Compression with Aligned Cross-View PriorsYongqi Zhai, Luyang Tang, Yi Ma, Rui Peng et al.ACM MM 2022 · 10 citations
- SASIC: Stereo Image Compression with Latent Shifts and Stereo AttentionMatthias Wödlinger, Jan Kotera, Jan Xu, Robert SablatnigCVPR 2022 · 24 citations
- Deep Homography for Efficient Stereo Image CompressionXin Deng, Wenzhe Yang, Ren Yang, Mai Xu et al.CVPR 2021
- DSIC: Deep Stereo Image CompressionJerry Liu, Shenlong Wang, Raquel UrtasunICCV 2019 · 50 citations
