MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image Restoration
Heng Li, Xingyuan Wang, Yang Fan, Yunan Zhang, Xiangping Wu, Qingcai Chen
摘要
Restoring degraded document image is essential for both improving visual quality and optimizing performance in downstream document analysis tasks. Although existing methods have demonstrated substantial improvements in restoration outcomes, they primarily address single-type degradation scenarios. Current approaches typically necessitate training multiple specialized models for specific degradation types or rely on explicit prior knowledge of degradation patterns to guide the training process. To overcome these limitations, we propose MMDIR, a multimodal instruction-driven framework designed for document image restoration under mixed and uncertain degradation conditions. By leveraging semantically structured instructions, MMDIR dynamically identifies present degradation types (blur, shadow, text watermark, and seal), while enhancing degradation-aware representation learning. Furthermore, we introduce a novel benchmark named MixedDoc comprising complex mixed degradations, where each image contains randomized combinations of the aforementioned types. This benchmark addresses a critical gap in existing datasets, which lack realistic multi-degradation samples and often overlook common obstructions such as seals and text watermarks. The effectiveness of our approach is thoroughly validated across both released public benchmarks and our newly proposed dataset. The dataset is available at https://github.com/xiaomore/MMDIR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing NetZinuo Li, Xuhang Chen, Chi-Man Pun, Xiaodong CunICCV 2023 · 被引用 66 次
- DocDiff: Document Enhancement via Residual Diffusion ModelsZongyuan Yang, Baolin Liu, Yongping Xiong, Lan Yi 等ACM MM 2023 · 被引用 55 次
- Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion ModelDian Zheng, Xiao-Ming Wu, Shuzhou Yang, Jian Zhang 等CVPR 2024 · 被引用 45 次
相关 Paper
- Visual-Instructed Degradation Diffusion for All-in-One Image RestorationWenyang Luo, Haina Qin, Zewen Chen, Libin Wang 等CVPR 2025
- DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksJiaxin Zhang, Dezhi Peng, Chongyu Liu, Peirong Zhang 等CVPR 2024
- PromptIR: Prompting for All-in-One Image RestorationVaishnav Potlapalli, Syed Waqas Zamir, Salman H. Khan, Fahad Shahbaz KhanNeurIPS 2023 · 被引用 386 次
- Dual-Level Prototype Learning for Composite Degraded Image RestorationZhongze Wang, Haitao Zhao, Lujian Yao, Jingchao Peng 等ICCV 2025
- Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics RecoveryZizheng Yang, Jie Huang, Jiahao Chang, Man Zhou 等CVPR 2023
