MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
Han Liu, Hengyu Man, Xingtao Wang, Wenrui Li, Debin Zhao
摘要
Recent advances in extreme image compression have revealed that mapping pixel data into highly compact latent representations can significantly improve coding efficiency. However, most existing methods compress images into 2-D latent spaces via convolutional neural networks (CNNs) or Swin Transformers, which tend to retain substantial spatial redundancy, thereby limiting overall compression performance. In this paper, we propose a novel Mixed RWKV-Transformer (MRT) architecture that encodes images into more compact 1-D latent representations by synergistically integrating the complementary strengths of linear-attentionbased RWKV and self-attention-based Transformer models. Specifically, MRT partitions each image into fixed-size windows, utilizing RWKV modules to capture global dependencies across windows and Transformer blocks to model local redundancies within each window. The hierarchical attention mechanism enables more efficient and compact representation learning in the 1-D domain. To further enhance compression efficiency, we introduce a dedicated RWKV Compression Model (RCM) tailored to the structure characteristics of the intermediate 1-D latent features in MRT. Extensive experiments on standard image compression benchmarks validate the effectiveness of our approach. The proposed MRT framework consistently achieves superior reconstruction quality at bitrates below 0.02 bits per pixel (bpp). Quantitative results based on the DISTS metric show that MRT significantly outperforms the state-of-the-art 2-D architecture GLC, achieving bitrate savings of 43.75%, 30.59% on the Kodak and CLIC2020 test datasets, respectively. All source code are available at https://github.com/luke1453lh/MRT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 被引用 675 次
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu 等CVPR 2022 · 被引用 346 次
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen 等NeurIPS 2024 · 被引用 331 次
- Lossy Image Compression with Conditional Diffusion ModelsRuihan Yang, Stephan MandtNeurIPS 2023 · 被引用 268 次
相关 Paper
- Linear Attention Modeling for Learned Image CompressionDonghui Feng, Zhengxue Cheng, Shen Wang, Ronghua Wu 等CVPR 2025
- Learned Image Compression with Mixed Transformer-CNN ArchitecturesJinming Liu, Heming Sun, Jiro KattoCVPR 2023
- Generative Latent Coding for Ultra-Low Bitrate Image CompressionZhaoyang Jia, Jiahao Li, Bin Li, Houqiang Li 等CVPR 2024
- Frequency-Aware Transformer for Learned Image CompressionHan Li, Shaohui Li, Wenrui Dai, Chenglin Li 等ICLR 2024 · 被引用 88 次
- The Devil Is in the Details: Window-based Attention for Image CompressionRenjie Zou, Chunfeng Song, Zhaoxiang ZhangCVPR 2022 · 被引用 260 次
