Multi-Modality Deep Network for Extreme Learned Image Compression
Xuhao Jiang, Weimin Tan, Tian Tan, Bo Yan, Liquan Shen
摘要
Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years , but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2x to 4x bitrates of ours.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual FidelityHagyeong Lee, Minkyu Kim, Jun-Hyuk Kim, Seungeon Kim 等ICML 2024 · 被引用 25 次
- Addressing Imbalance for Class Incremental Learning in Medical Image ClassificationXuze Hao, Wenqian Ni, Xuhao Jiang, Weimin Tan 等ACM MM 2024 · 被引用 7 次
- msLPCC: A Multimodal-Driven Scalable Framework for Deep LiDAR Point Cloud CompressionMiaohui Wang, Runnan Huang, Hengjin Dong, Di Lin 等AAAI 2024 · 被引用 7 次
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image CompressionHan Liu, Hengyu Man, Xingtao Wang, Wenrui Li 等AAAI 2026 · 被引用 2 次
- DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionYouneng Bao, Yulong Cheng, Yiping Liu, Yichen Yang 等AAAI 2026
它引用的顶会 Paper8
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte 等ICCV 2019 · 被引用 648 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- Enhanced Invertible Encoding for Learned Image CompressionYueqi Xie, Ka Leong Cheng, Qifeng ChenACM MM 2021 · 被引用 195 次
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 被引用 143 次
相关 Paper
- Towards Unified Human Perception and Machine Understanding: Token Flow Guided Compression FrameworkLi Xu, Yingfu Zhang, Kepeng Xu, Gang He 等CVPR 2026
- Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine PerceptionYudong Mao, Peilin Chen, Shurun Wang, Shiqi Wang 等ACM MM 2023 · 被引用 5 次
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma 等ACM MM 2021 · 被引用 24 次
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An 等ICLR 2025
- Distributed Image Compression with Multimodal Side Information at Extremely Low BitratesGuojun Xu, Mingyang Zhang, Jianwen Xiang, Cheng Tan 等CVPR 2026
