Frequency-Aware Transformer for Learned Image Compression
Han Li, Shaohui Li, Wenrui Dai, Chenglin Li, Junni Zou, Hongkai Xiong
摘要
Learned image compression (LIC) has gained traction as an effective solution for image storage and transmission in recent years. However, existing LIC methods are redundant in latent representation due to limitations in capturing anisotropic frequency components and preserving directional details. To overcome these challenges, we propose a novel frequency-aware transformer (FAT) block that for the first time achieves multiscale directional ananlysis for LIC. The FAT block comprises frequency-decomposition window attention (FDWA) modules to capture multiscale and directional frequency components of natural images. Additionally, we introduce frequency-modulation feed-forward network (FMFFN) to adaptively modulate different frequency components, improving rate-distortion performance. Furthermore, we present a transformer-based channel-wise autoregressive (T-CA) model that effectively exploits channel dependencies. Experiments show that our method achieves state-of-the-art rate-distortion performance compared to existing LIC methods, and evidently outperforms latest standardized codec VTM-12.1 by 14.5%, 15.1%, 13.0% in BD-rate on the Kodak, Tecnick, and CLIC datasets. Code will be released at https://github.com/qingshi9974/ ICLR2024-FTIC
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Improving Diffusion Models for Inverse Problems Using Optimal Posterior CovarianceXinyu Peng, Ziyang Zheng, Wenrui Dai, Nuoqian Xiao 等ICML 2024 · 被引用 47 次
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and GenerationHan Li, Xinyu Peng, Yaoming Wang, Zelin Peng 等CVPR 2026 · 被引用 47 次
- Causal Context Adjustment Loss for Learned Image CompressionMinghao Han, Shiyin Jiang, Shengxi Li, Xin Deng 等NeurIPS 2024 · 被引用 31 次
- Learned Image Compression with Hierarchical Progressive Context ModelingYuqi Li, Haotian Zhang, Li Li, Dong LiuICCV 2025 · 被引用 8 次
- Knowledge Distillation for Learned Image CompressionYunuo Chen, Zezheng Lyu, Bing He, Ning Cao 等ICCV 2025 · 被引用 7 次
它引用的顶会 Paper16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu 等NeurIPS 2021 · 被引用 798 次
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 被引用 653 次
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma 等CVPR 2022 · 被引用 363 次
相关 Paper
- Learned Image Compression via Sparse Attention and Adaptive FrequencyHuidong Ma, Xinyan Shi, Hui Sun, Xiaofei Yue 等CVPR 2026
- Learned Image Compression with Mixed Transformer-CNN ArchitecturesJinming Liu, Heming Sun, Jiro KattoCVPR 2023
- Linear Attention Modeling for Learned Image CompressionDonghui Feng, Zhengxue Cheng, Shen Wang, Ronghua Wu 等CVPR 2025
- Few-Shot Domain Adaptation for Learned Image CompressionTianyu Zhang, Haotian Zhang, Yuqi Li, Li Li 等AAAI 2025 · 被引用 2 次
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
