PICD: Versatile Perceptual Image Compression with Diffusion Rendering
Tongda Xu, Jiahao Li, Bin Li, Yan Wang, Ya-Qin Zhang, Yan Lu
Abstract
Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- One-Step Diffusion-Based Image Compression with Semantic DistillationNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.NeurIPS 2025 · 28 citations
- Benchmarking and Enhancing VLM for Compressed Image UnderstandingZifu Zhang, Tongda Xu, Siqi Li, Shengxi Li et al.ICML 2026 · 2 citations
- DLF: Extreme Image Compression with Dual-Generative Latent FusionNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.ICCV 2025 · 2 citations
- BADiff: Bandwidth Adaptive Diffusion ModelXi Zhang, Hanwei Zhu, Yan Zhong, Jiamang Wang et al.NeurIPS 2025 · 1 citation
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An et al.ICLR 2025
- CADC: Content Adaptive Diffusion-Based Generative Image CompressionXihua Sheng, Lingyu Zhu, Tianyu Zhang, Dong Liu et al.CVPR 2026 · 2 citations
- Single-step Diffusion-based Video Coding with Semantic-Temporal GuidanceNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li et al.CVPR 2026 · 12 citations
- Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual FidelityHagyeong Lee, Minkyu Kim, Jun-Hyuk Kim, Seungeon Kim et al.ICML 2024 · 25 citations
- CoD: A Diffusion Foundation Model for Image CompressionZhaoyang Jia, Zihan Zheng, Naifu Xue, Jiahao Li et al.CVPR 2026 · 9 citations
