PICD: Versatile Perceptual Image Compression with Diffusion Rendering
Tongda Xu, Jiahao Li, Bin Li, Yan Wang, Ya-Qin Zhang, Yan Lu
摘要
Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- One-Step Diffusion-Based Image Compression with Semantic DistillationNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li 等NeurIPS 2025 · 被引用 28 次
- Benchmarking and Enhancing VLM for Compressed Image UnderstandingZifu Zhang, Tongda Xu, Siqi Li, Shengxi Li 等ICML 2026 · 被引用 2 次
- DLF: Extreme Image Compression with Dual-Generative Latent FusionNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li 等ICCV 2025 · 被引用 2 次
- BADiff: Bandwidth Adaptive Diffusion ModelXi Zhang, Hanwei Zhu, Yan Zhong, Jiamang Wang 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- DiffPC: Diffusion-based High Perceptual Fidelity Image Compression with Semantic RefinementYichong Xia, Yimin Zhou, Jinpeng Wang, Baoyi An 等ICLR 2025
- CADC: Content Adaptive Diffusion-Based Generative Image CompressionXihua Sheng, Lingyu Zhu, Tianyu Zhang, Dong Liu 等CVPR 2026 · 被引用 2 次
- Single-step Diffusion-based Video Coding with Semantic-Temporal GuidanceNaifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li 等CVPR 2026 · 被引用 12 次
- Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual FidelityHagyeong Lee, Minkyu Kim, Jun-Hyuk Kim, Seungeon Kim 等ICML 2024 · 被引用 25 次
- CoD: A Diffusion Foundation Model for Image CompressionZhaoyang Jia, Zihan Zheng, Naifu Xue, Jiahao Li 等CVPR 2026 · 被引用 9 次
