Generative Latent Coding for Ultra-Low Bitrate Image Compression
Zhaoyang Jia, Jiahao Li, Bin Li, Houqiang Li, Yan Lu
Abstract
Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial redundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the pixel-space distortion may not align with human perception. To address this issue, we introduce a Generative Latent Coding (GLC) architecture, which performs transform coding in the latent space of a generative vector-quantized variational auto-encoder (VQ-VAE), instead of in the pixel space. The generative latent space is characterized by greater sparsity, richer semantic and better alignment with human perception, rendering it advantageous for achieving high-realism and high-fidelity compression. Additionally, we introduce a categorical hyper module to reduce the bit cost of hyper-information, and a code-prediction-based supervision to enhance the semantic consistency. Experiments demonstrate that our GLC maintains high visual quality with less than 0.04 bpp on natural images and less than 0.01 bpp on facial images. On the CLIC2020 test set, we achieve the same FID as MS-ILLM with 45% fewer bits. Furthermore, the powerful generative latent space enables various applications built on our GLC pipeline, such as image restoration and style transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d99c4a94-9787-411e-bd22-cec86efdb036Cited by top-tier papers26
- One-Step Diffusion-Based Image Compression with Semantic DistillationNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.NeurIPS 2025 · 28 citations
- Generative Neural Video Compression via Video Diffusion PriorQi Mao, Hao Cheng, Tinghan Yang, Libiao Jin et al.CVPR 2026 · 18 citations
- Single-step Diffusion-based Video Coding with Semantic-Temporal GuidanceNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.CVPR 2026 · 12 citations
- Conditional Latent Coding with Learnable Synthesized Reference for Deep Image CompressionSiqi Wu, Yinda Chen, Dong Liu, Zhihai HeAAAI 2025 · 9 citations
- CoD: A Diffusion Foundation Model for Image CompressionZhaoyang Jia, Zihan Zheng, Naifu Xue, Jiahao Li et al.CVPR 2026 · 9 citations
Builds on19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat et al.CVPR 2022 · 3,348 citations
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 675 citations
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
Related papers
- Improving Statistical Fidelity for Neural Image Compression with Implicit Local Likelihood ModelsMatthew J. Muckley, Alaaeldin El-Nouby, Karen Ullrich, Hervé Jégou et al.ICML 2023 · 114 citations
- Decouple Distortion from Perception: Region Adaptive Diffusion for Extreme-low Bitrate Perception Image CompressionJinchang Xu, Shaokang Wang, Jintao Chen, Zhe Li et al.CVPR 2025
- DLF: Extreme Image Compression with Dual-Generative Latent FusionNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.ICCV 2025 · 2 citations
- Efficient Learned Image Compression without Entropy CodingHao Cao, Wenqi Guo, Zhijin Qin, Jungong HanICML 2026
- SSCL: Adversarially Guided Image Compression via Semantic and Spectral Consistency LearningWei Jiang, Yongqi Zhai, Jiayu Yang, Bohao Feng et al.AAAI 2026
