Entroformer: A Transformer-based Entropy Model for Learned Image Compression
Yichen Qian, Xiuyu Sun, Ming Lin, Zhiyu Tan, Rong Jin
Abstract
One critical component in lossy deep image compression is the entropy model, which predicts the probability distribution of the quantized latent representation in the encoding and decoding modules. Previous works build entropy models upon convolutional neural networks which are inefficient in capturing global dependencies. In this work, we propose a novel transformer-based entropy model, termed Entroformer, to capture long-range dependencies in probability distribution estimation effectively and efficiently. Different from vision transformers in image classification, the Entroformer is highly optimized for image compression, including a top-k self-attention and a diamond relative position encoding. Meanwhile, we further expand this architecture with a parallel bidirectional context model to speed up the decoding process. The experiments show that the Entroformer achieves state-of-the-art performance on image compression while being time-efficient. Code is available at https://github.com/damo-cv/entroformer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d87cf03-3aff-433d-b15b-7d19b16e5e6cCited by top-tier papers43
- CDTrans: Cross-domain Transformer for Unsupervised Domain AdaptationTongkun Xu, Weihua Chen, Pichao Wang, Fan Wang et al.ICLR 2022 · 293 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
- MLIC: Multi-Reference Entropy Model for Learned Image CompressionWei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning et al.ACM MM 2023 · 117 citations
- Frequency-Aware Transformer for Learned Image CompressionHan Li, Shaohui Li, Wenrui Dai, Chenglin Li et al.ICLR 2024 · 88 citations
- Causal Context Adjustment Loss for Learned Image CompressionMinghao Han, Shiyin Jiang, Shengxi Li, Xin Deng et al.NeurIPS 2024 · 31 citations
Builds on5
- Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionYueyu Hu, Wenhan Yang, Jiaying LiuAAAI 2020 · 143 citations
- Learning Accurate Entropy Model with Global Reference for Image CompressionYichen Qian, Zhiyu Tan, Xiuyu Sun, Ming Lin et al.ICLR 2021 · 93 citations
- Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure SpaceYaohua Wang, Yaobin Zhang, Fangyi Zhang, Senzhang Wang et al.ICLR 2022 · 38 citations
- Checkerboard Context Model for Efficient Learned Image CompressionDailan He, Yaoyan Zheng, Baocheng Sun, Yan Wang et al.CVPR 2021
- Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention ModulesZhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro KattoCVPR 2020
Related papers
- Joint Global and Local Hierarchical Priors for Learned Image CompressionJun-Hyuk Kim, Byeongho Heo, Jong-Seok LeeCVPR 2022 · 98 citations
- Laplacian-guided Entropy Model in Neural Codec with Blur-dissipated SynthesisAtefeh Khoshkhahtinat, Ali Zafari, Piyush M. Mehta, Nasser M. NasrabadiCVPR 2024
- Learned Image Compression with Dictionary-based Entropy ModelJingbo Lu, Leheng Zhang, Xingyu Zhou, Mu Li et al.CVPR 2025
- MIMT: Masked Image Modeling Transformer for Video CompressionJinxi Xiang, Kuan Tian, Jun ZhangICLR 2023
- Learned Image Compression with Mixed Transformer-CNN ArchitecturesJinming Liu, Heming Sun, Jiro KattoCVPR 2023
