CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
Xinjie Zhang, Shenyuan Gao, Zhening Liu, Jiawei Shao, Xingtong Ge, Dailan He, Tongda Xu, Yan Wang, Jun Zhang
摘要
Existing learning-based stereo image codec adopt sophisticated transformation with simple entropy models derived from single image codecs to encode latent representations. However, those entropy models struggle to effectively capture the spatial-disparity characteristics inherent in stereo images, which leads to suboptimal rate-distortion results. In this paper, we propose a stereo image compression framework, named CAMSIC. CAMSIC independently transforms each image to latent representation and employs a powerful decoder-free Transformer entropy model to capture both spatial and disparity dependencies, by introducing a novel content-aware masked image modeling (MIM) technique. Our content-aware MIM facilitates efficient bidirectional interaction between prior information and estimated tokens, which naturally obviates the need for an extra Transformer decoder. Experiments show that our stereo image codec achieves state-of-the-art rate-distortion performance on two stereo image datasets Cityscapes and InStereo2K with fast encoding and decoding speed. * This work was partially performed when Xinjie Zhang was an Intern at SenseTime.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy ModelShiyu Qin, XINJIE ZHANG, Zhening Liu, Jinpeng Wang 等CVPR 2026
- Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image CompressionHaotian Zhang, Feiyue Long, Yixin Yu, Jian Xue 等CVPR 2026
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li 等CVPR 2026
- Distributed Image Compression with Multimodal Side Information at Extremely Low BitratesGuojun Xu, Mingyang Zhang, Jianwen Xiang, Cheng Tan 等CVPR 2026
- SDiD:Shared diffusion prior for efficient distributed stereo image compressionYichong Xia, Yimin Zhou, Zongyu Li, Shiyu Qin 等ICML 2026
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- Masked Autoencoders As Spatiotemporal LearnersChristoph Feichtenhofer, Haoqi Fan, Yanghao Li, Kaiming HeNeurIPS 2022 · 被引用 690 次
相关 Paper
- Deep Stereo Image Compression via Bi-directional CodingJianjun Lei, Xiangrui Liu, Bo Peng, Dengchao Jin 等CVPR 2022 · 被引用 18 次
- Disparity-based Stereo Image Compression with Aligned Cross-View PriorsYongqi Zhai, Luyang Tang, Yi Ma, Rui Peng 等ACM MM 2022 · 被引用 10 次
- SASIC: Stereo Image Compression with Latent Shifts and Stereo AttentionMatthias Wödlinger, Jan Kotera, Jan Xu, Robert SablatnigCVPR 2022 · 被引用 24 次
- Deep Homography for Efficient Stereo Image CompressionXin Deng, Wenzhe Yang, Ren Yang, Mai Xu 等CVPR 2021
- DSIC: Deep Stereo Image CompressionJerry Liu, Shenlong Wang, Raquel UrtasunICCV 2019 · 被引用 50 次
