Multi-Modality Deep Network for Extreme Learned Image Compression
Xuhao Jiang, Weimin Tan, Tian Tan, Bo Yan, Liquan Shen
Abstract
Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual FidelityHagyeong Lee, Minkyu Kim, Jun-Hyuk Kim, Seungeon Kim et al.ICML 2024 · 25 citations
- Addressing Imbalance for Class Incremental Learning in Medical Image ClassificationXuze Hao, Wenqian Ni, Xuhao Jiang, Weimin Tan et al.ACM MM 2024 · 7 citations
- msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionMiaohui Wang, Runnan Huang, Hengjin Dong, Di Lin et al.AAAI 2024 · 7 citations
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image CompressionHan Liu, Hengyu Man, Xingtao Wang, Wenrui Li et al.AAAI 2026 · 2 citations
- DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionYouneng Bao, Yulong Cheng, Yiping Liu, Yichen Yang et al.AAAI 2026
Builds on8
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Enhanced Invertible Encoding for Learned Image CompressionYueqi Xie, Ka Leong Cheng, Qifeng ChenACM MM 2021 · 195 citations
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 143 citations
Related papers
- Towards Unified Human Perception and Machine Understanding: Token Flow Guided Compression FrameworkLi Xu, Yingfu Zhang, Kepeng Xu, Gang He et al.CVPR 2026
- Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine PerceptionYudong Mao, Peilin Chen, Shurun Wang, Shiqi Wang et al.ACM MM 2023 · 5 citations
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma et al.ACM MM 2021 · 24 citations
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An et al.ICLR 2025
- Distributed Image Compression with Multimodal Side Information at Extremely Low BitratesGuojun Xu, Mingyang Zhang, Jianwen Xiang, Cheng Tan et al.CVPR 2026
