Frequency-Aware Transformer for Learned Image Compression
Han Li, Shaohui Li, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong
Abstract
Learned image compression (LIC) has gained traction as an effective solution for image storage and transmission in recent years. However, existing LIC methods are redundant in latent representation due to limitations in capturing anisotropic frequency components and preserving directional details. To overcome these challenges, we propose a novel frequency-aware transformer (FAT) block that for the first time achieves multiscale directional ananlysis for LIC. The FAT block comprises frequency-decomposition window attention (FDWA) modules to capture multiscale and directional frequency components of natural images. Additionally, we introduce frequency-modulation feed-forward network (FMFFN) to adaptively modulate different frequency components, improving rate-distortion performance. Furthermore, we present a transformer-based channel-wise autoregressive (T-CA) model that effectively exploits channel dependencies. Experiments show that our method achieves state-of-the-art rate-distortion performance compared to existing LIC methods, and evidently outperforms latest standardized codec VTM-12.1 by 14.5%, 15.1%, 13.0% in BD-rate on the Kodak, Tecnick, and CLIC datasets. Code will be released at https://github.com/qingshi9974/ ICLR2024-FTIC
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 032c8f72-9ed5-4b56-bae1-943313b526f4Cited by top-tier papers27
- Improving Diffusion Models for Inverse Problems Using Optimal Posterior CovarianceXinyu Peng, Ziyang Zheng, Wenrui Dai, Nuoqian Xiao et al.ICML 2024 · 47 citations
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and GenerationHan Li, Xinyu Peng, Yaoming Wang, Zelin Peng et al.CVPR 2026 · 47 citations
- Causal Context Adjustment Loss for Learned Image CompressionMinghao Han, Shiyin Jiang, Shengxi Li, Xin Deng et al.NeurIPS 2024 · 31 citations
- Learned Image Compression with Hierarchical Progressive Context ModelingYuqi Li, Haotian Zhang, Li Li, Dong LiuICCV 2025 · 8 citations
- Knowledge Distillation for Learned Image CompressionYunuo Chen, Zezheng Lyu, Bing He, Ning Cao et al.ICCV 2025 · 7 citations
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
Related papers
- Learned Image Compression via Sparse Attention and Adaptive FrequencyHuidong Ma, Xinyan Shi, Hui Sun, Xiaofei Yue et al.CVPR 2026
- Learned Image Compression with Mixed Transformer-CNN ArchitecturesJinming Liu, Heming Sun, Jiro KattoCVPR 2023
- Linear Attention Modeling for Learned Image CompressionDonghui Feng, Zhengxue Cheng, Shen Wang, Ronghua Wu et al.CVPR 2025
- Few-Shot Domain Adaptation for Learned Image CompressionTianyu Zhang, Haotian Zhang, Yuqi Li, Li Li et al.AAAI 2025 · 2 citations
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
