Linear Attention Modeling for Learned Image Compression
Donghui Feng, Zhengxue Cheng, Shen Wang, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song
Abstract
Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural networkbased transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKVbased Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https: //github.com/sjtu-medialab/RwkvCompress .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 496a3af1-8f2c-4fc4-b1c7-1afcfe5bea37Cited by top-tier papers12
- Content-Aware Mamba for Learned Image CompressionYunuo Chen, Zezheng Lyu, Bing He, Hongwei Hu et al.ICLR 2026 · 5 citations
- GIViC: Generative Implicit Video CompressionGe Gao, Siyue Teng, Tianhao Peng, Fan Zhang et al.ICCV 2025 · 4 citations
- TaCo: A Benchmark for Lossless and Lossy Codecs of Heterogeneous Tactile DataZhengxue Cheng, Yan Zhao, Keyu Wang, Hengdi Zhang et al.ICLR 2026 · 3 citations
- OmniZip: Learning a Unified and Lightweight Lossless Compressor for Multi-Modal DataYan Zhao, Zhengxue Cheng, Junxuan Zhang, Dajiang Zhou et al.CVPR 2026 · 2 citations
- CADC: Content Adaptive Diffusion-Based Generative Image CompressionXihua Sheng, Lingyu Zhu, Tianyu Zhang, Dong Liu et al.CVPR 2026 · 2 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- The Devil Is in the Details: Window-based Attention for Image CompressionRenjie Zou, Chunfeng Song, Zhaoxiang ZhangCVPR 2022 · 260 citations
Related papers
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image CompressionHan Liu, Hengyu Man, Xingtao Wang, Wenrui Li et al.AAAI 2026 · 2 citations
- Learned Image Compression with Mixed Transformer-CNN ArchitecturesJinming Liu, Heming Sun, Jiro KattoCVPR 2023
- MLIC: Multi-Reference Entropy Model for Learned Image CompressionWei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning et al.ACM MM 2023 · 117 citations
- Learned Image Compression via Sparse Attention and Adaptive FrequencyHuidong Ma, Xinyan Shi, Hui Sun, Xiaofei Yue et al.CVPR 2026
- Another Way to the Top: Exploit Contextual Clustering in Learned Image CodingYichi Zhang, Zhihao Duan, Ming Lu, Dandan Ding et al.AAAI 2024 · 13 citations
