MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image Restoration
Heng Li, Xingyuan Wang, Yang Fan, Yunan Zhang, Xiangping Wu, Qingcai Chen
Abstract
Restoring degraded document image is essential for both improving visual quality and optimizing performance in downstream document analysis tasks. Although existing methods have demonstrated substantial improvements in restoration outcomes, they primarily address single-type degradation scenarios. Current approaches typically necessitate training multiple specialized models for specific degradation types or rely on explicit prior knowledge of degradation patterns to guide the training process. To overcome these limitations, we propose MMDIR, a multimodal instruction-driven framework designed for document image restoration under mixed and uncertain degradation conditions. By leveraging semantically structured instructions, MMDIR dynamically identifies present degradation types (blur, shadow, text watermark, and seal), while enhancing degradation-aware representation learning. Furthermore, we introduce a novel benchmark named MixedDoc comprising complex mixed degradations, where each image contains randomized combinations of the aforementioned types. This benchmark addresses a critical gap in existing datasets, which lack realistic multi-degradation samples and often overlook common obstructions such as seals and text watermarks. The effectiveness of our approach is thoroughly validated across both released public benchmarks and our newly proposed dataset. The dataset is available at https://github.com/xiaomore/MMDIR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b21ffa2b-aa8f-4283-8791-9cadb5e3d185Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing NetZinuo Li, Xuhang Chen, Chi-Man Pun, Xiaodong CunICCV 2023 · 66 citations
- DocDiff: Document Enhancement via Residual Diffusion ModelsZongyuan Yang, Baolin Liu, Yongping Xiong, Lan Yi et al.ACM MM 2023 · 55 citations
- Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion ModelDian Zheng, Xiao-Ming Wu, Shuzhou Yang, Jian Zhang et al.CVPR 2024 · 45 citations
Related papers
- Visual-Instructed Degradation Diffusion for All-in-One Image RestorationWenyang Luo, Haina Qin, Zewen Chen, Libin Wang et al.CVPR 2025
- DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksJiaxin Zhang, Dezhi Peng, Chongyu Liu, Peirong Zhang et al.CVPR 2024
- PromptIR: Prompting for All-in-One Image RestorationVaishnav Potlapalli, Syed Waqas Zamir, Salman H. Khan, Fahad Shahbaz KhanNeurIPS 2023 · 386 citations
- Dual-Level Prototype Learning for Composite Degraded Image RestorationZhongze Wang, Haitao Zhao, Lujian Yao, Jingchao Peng et al.ICCV 2025
- Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics RecoveryZizheng Yang, Jie Huang, Jiahao Chang, Man Zhou et al.CVPR 2023
