DocDiff: Document Enhancement via Residual Diffusion Models
Zongyuan Yang, Baolin Liu, Yongping Xiong, Lan Yi, Guibin Wu, Xiaojun Tang, Ziqi Liu, Junjie Zhou, Xing Zhang
Abstract
Removing degradation from document images not only improves their visual quality and readability, but also enhances the performance of numerous automated document analysis and recognition tasks. However, existing regression-based methods optimized for pixel-level distortion reduction tend to suffer from significant loss of high-frequency information, leading to distorted and blurred text edges. To compensate for this major deficiency, we propose DocDiff, the first diffusion-based framework specifically designed for diverse challenging document enhancement problems, including document deblurring, denoising, and removal of watermarks and seals. DocDiff consists of two modules: the Coarse Predictor (CP), which is responsible for recovering the primary low-frequency content, and the High-Frequency Residual Refinement (HRR) module, which adopts the diffusion models to predict the residual (high-frequency information, including text edges), between the ground-truth and the CP-predicted image. DocDiff is a compact and computationally efficient model that benefits from a well-designed network architecture, an optimized training loss objective, and a deterministic sampling process with short time steps. Extensive experiments demonstrate that DocDiff achieves state-of-the-art (SOTA) performance on multiple benchmark datasets, and can significantly enhance the readability and recognizability of degraded document images. Furthermore, our proposed HRR module in pre-trained DocDiff is plug-and-play and ready-to-use, with only 4.17M parameters. It greatly sharpens the text edges generated by SOTA deblurring methods without additional joint training. Available codes: https://github.com/Royalvice/DocDiff https://github.com/Royalvice/DocDiff.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Wavelet-based Fourier Information Interaction with Frequency Diffusion Adjustment for Underwater Image RestorationChen Zhao, Weiling Cai, Chenyu Dong, Chengwei HuCVPR 2024 · 116 citations
- PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-ResolutionZuoyan Zhao, Hui Xue, Pengfei Fang, Shipeng ZhuACM MM 2024 · 16 citations
- Predicting the Original Appearance of Damaged Historical DocumentsZhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang et al.AAAI 2025 · 8 citations
- Reproducing the Past: A Dataset for Benchmarking Inscription RestorationShipeng Zhu, Hui Xue, Na Nie, Chenjie Zhu et al.ACM MM 2024 · 4 citations
- EpiAgent: An Agent-Centric System for Ancient Inscription RestorationShipeng Zhu, Ang Chen, Na Nie, Pengfei Fang et al.CVPR 2026 · 2 citations
Builds on10
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Palette: Image-to-Image Diffusion ModelsChitwan Saharia, William Chan, Huiwen Chang, Chris A. Lee et al.SIGGRAPH 2022 · 1,638 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionFangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang et al.ACM MM 2025 · 1 citation
- MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image RestorationHeng Li, Xingyuan Wang, Yang Fan, Yunan Zhang et al.CVPR 2026
- End-to-End Unsupervised Document Image Blind DenoisingMehrdad J. Gangeh, Marcin Plata, Hamid R. Motahari Nezhad, Nigel P. DuffyICCV 2021 · 14 citations
- Document Registration: Towards Automated Labeling of Pixel-Level Alignment Between Warped-Flat DocumentsWeiguang Zhang, Qiufeng Wang, Kaizhu Huang, Xiaowei Huang et al.ACM MM 2024 · 1 citation
- DERO: Diffusion-Model-Erasure Robust WatermarkingHan Fang, Kejiang Chen, Yupeng Qiu, Zehua Ma et al.ACM MM 2024 · 6 citations
