Learned Image Compression with Mixed Transformer-CNN Architectures
Jinming Liu, Heming Sun, Jiro Katto
Abstract
Learned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existing LIC methods are Convolutional Neural Networks-based (CNN-based) or Transformer-based, which have different advantages. Exploiting both advantages is a point worth exploring, which has two challenges: 1) how to effectively fuse the two methods? 2) how to achieve higher performance with a suitable complexity? In this paper, we propose an efficient parallel Transformer-CNN Mixture (TCM) block with a controllable complexity to incorporate the local modeling ability of CNN and the non-local modeling ability of transformers to improve the overall architecture of image compression models. Besides, inspired by the recent progress of entropy estimation models and attention modules, we propose a channel-wise entropy model with parameter-efficient swin-transformer-based attention (SWAtten) modules by using channel squeezing. Experimental results demonstrate our proposed method achieves state-of-the-art rate-distortion performances on three different resolution datasets (i.e., Kodak, Tecnick, CLIC Professional Validation) compared to existing LIC methods. The code is at https://github.com/jmliu206/ LIC_TCM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a063ef9a-427b-46bc-ac85-22ccb42958adCited by top-tier papers76
- Frequency-Aware Transformer for Learned Image CompressionHan Li, Shaohui Li, Wenrui Dai, Chenglin Li et al.ICLR 2024 · 88 citations
- Compression with Bayesian Implicit Neural RepresentationsZongyu Guo, Gergely Flamich, Jiajun He, Zhibo Chen et al.NeurIPS 2023 · 38 citations
- Idempotence and Perceptual Image CompressionTongda Xu, Ziran Zhu, Dailan He, Yanghao Li et al.ICLR 2024 · 33 citations
- Causal Context Adjustment Loss for Learned Image CompressionMinghao Han, Shiyin Jiang, Shengxi Li, Xin Deng et al.NeurIPS 2024 · 31 citations
- One-Step Diffusion-Based Image Compression with Semantic DistillationNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.NeurIPS 2025 · 28 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- The Devil Is in the Details: Window-based Attention for Image CompressionRenjie Zou, Chunfeng Song, Zhaoxiang ZhangCVPR 2022 · 260 citations
Related papers
- Joint Global and Local Hierarchical Priors for Learned Image CompressionJun-Hyuk Kim, Byeongho Heo, Jong-Seok LeeCVPR 2022 · 98 citations
- Linear Attention Modeling for Learned Image CompressionDonghui Feng, Zhengxue Cheng, Shen Wang, Ronghua Wu et al.CVPR 2025
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image CompressionHan Liu, Hengyu Man, Xingtao Wang, Wenrui Li et al.AAAI 2026 · 2 citations
- MLIC: Multi-Reference Entropy Model for Learned Image CompressionWei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning et al.ACM MM 2023 · 117 citations
- Entroformer: A Transformer-based Entropy Model for Learned Image CompressionYichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan et al.ICLR 2022 · 194 citations
