Two-Stage Octave Residual Network for End-to-End Image Compression
Fangdong Chen, Yumeng Xu, Li Wang
Abstract
Octave Convolution (OctConv) is a generic convolutional unit that has already achieved good performances in many computer vision tasks. Recent studies also have shown the potential of applying the OctConv in end-to-end image compression. However, considering the characteristic of image compression task, current works of OctConv may limit the performance of the image compression network due to the loss of spatial information caused by the sampling operations of inter-frequency communication. Besides, the correlation between multi-frequency latents produced by OctConv is not utilized in current architectures. In this paper, to address these problems, we propose a novel Two-stage Octave Residual (ToRes) block which strips the sampling operation from OctConv to strengthen the capability of preserving useful information. Moreover, to capture the redundancy between the multi-frequency latents, a context transfer module is designed. The results show that both ToRes block and the incorporation of context transfer module help to improve the Rate-Distortion performance, and the combination of these two strategies makes our model achieve the state-of-the-art performance and outperform the latest compression standard Versatile Video Coding (VVC) in terms of both PSNR and MS-SSIM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02ca42f5-8a57-4b8c-9158-39c91e3959f2Cited by top-tier papers12
- MLIC: Multi-Reference Entropy Model for Learned Image CompressionWei Jiang, Jiayu Yang, Yongqi Zhai, Peirong Ning et al.ACM MM 2023 · 117 citations
- Another Way to the Top: Exploit Contextual Clustering in Learned Image CodingYichi Zhang, Zhihao Duan, Ming Lu, Dandan Ding et al.AAAI 2024 · 13 citations
- Knowledge Distillation for Learned Image CompressionYunuo Chen, Zezheng Lyu, Bing He, Ning Cao et al.ICCV 2025 · 7 citations
- Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP SupervisionKangsheng Yin, Quan Liu, Xuelin Shen, Yulin He et al.AAAI 2025 · 6 citations
- Content-Aware Mamba for Learned Image CompressionYunuo Chen, Zezheng Lyu, Bing He, Hongwei Hu et al.ICLR 2026 · 5 citations
Builds on3
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan et al.ICCV 2019 · 665 citations
- A Spatial RNN Codec for End-to-End Image CompressionChaoyi Lin, Jiabao Yao, Fangdong Chen, Li WangCVPR 2020
- Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention ModulesZhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro KattoCVPR 2020
Related papers
- Learned Bi-Resolution Image Coding using Generalized Octave ConvolutionsMohammad Akbari, Jie Liang, Jingning Han, Chengjie TuAAAI 2021 · 21 citations
- Enhanced Invertible Encoding for Learned Image CompressionYueqi Xie, Ka Leong Cheng, Qifeng ChenACM MM 2021 · 195 citations
- Neural Image Compression via Attentional Multi-scale Back Projection and Frequency DecompositionGe Gao, Pei You, Rong Pan, Shunyuan Han et al.ICCV 2021 · 97 citations
- FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context ModelShiyu Qin, Yongkang Lu, Yimin Zhou, Jiawei Li et al.CVPR 2026
- Dual-Octave Convolution for Accelerated Parallel MR Image ReconstructionChun-Mei Feng, Zhanyuan Yang, Geng Chen, Yong Xu et al.AAAI 2021 · 31 citations
