End-to-End RGB-D Image Compression via Exploiting Channel-Modality Redundancy
Huiming Zheng, Wei Gao
摘要
As a kind of 3D data, RGB-D images have been extensively used in object tracking, 3D reconstruction, remote sensing mapping, and other tasks. In the realm of computer vision, the significance of RGB-D images is progressively growing. However, the existing learning-based image compression methods usually process RGB images and depth images separately, which cannot entirely exploit the redundant information between the modalities, limiting the further improvement of the Rate-Distortion performance. With the goal of overcoming the defect, in this paper, we propose a learning-based dual-branch RGB-D image compression framework. Compared with traditional RGB domain compression scheme, a YUV domain compression scheme is presented for spatial redundancy removal. In addition, Intra-Modality Attention (IMA) and Cross-Modality Attention (CMA) are introduced for modal redundancy removal. For the sake of benefiting from cross-modal prior information, Context Prediction Module (CPM) and Context Fusion Module (CFM) are raised in the conditional entropy model which makes the context probability prediction more accurate. The experimental results demonstrate our method outperforms existing image compression methods in two RGB-D image datasets. Compared with BPG, our proposed framework can achieve up to 15% bit rate saving for RGB images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- The Devil Is in the Details: Window-based Attention for Image CompressionRenjie Zou, Chunfeng Song, Zhaoxiang ZhangCVPR 2022 · 被引用 260 次
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 被引用 143 次
- MLIC: Multi-Reference Entropy Model for Learned Image CompressionWei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning 等ACM MM 2023 · 被引用 117 次
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang 等ACM MM 2020 · 被引用 53 次
相关 Paper
- Learning based Multi-modality Image and Video CompressionGuo Lu, Tianxiong Zhong, Jing Geng, Qiang Hu 等CVPR 2022 · 被引用 26 次
- Visual Redundancy Removal of Composite Images via Multimodal LearningWuyuan Xie, Shukang Wang, Rong Zhang, Miaohui WangACM MM 2023 · 被引用 1 次
- Spatial-Temporal Context Model for Remote Sensing Imagery CompressionJinxiao Zhang, Runmin Dong, Juepeng Zheng, Mengxuan Chen 等ACM MM 2024 · 被引用 6 次
- Learning Selective Self-Mutual Attention for RGB-D Saliency DetectionNian Liu, Ni Zhang, Junwei HanCVPR 2020
- Low-Latency Neural LiDAR Compression with 2D Context ModelsRui Song, Yan Wang, Tongda Xu, Zhening Liu 等ICLR 2026
