DocRes: A Generalist Model Toward Unifying Document Image Restoration Tasks
Jiaxin Zhang, Dezhi Peng, Chongyu Liu, Peirong Zhang, Lianwen Jin
摘要
Document image restoration is a crucial aspect of Document AI systems, as the quality of document images significantly influences the overall performance. Prevailing methods address distinct restoration tasks independently, leading to intricate systems and the incapability to harness the potential synergies of multi-task learning. To overcome this challenge, we propose DocRes, a generalist model that unifies five document image restoration tasks including dewarping, deshadowing, appearance enhancement, deblurring, and binarization. To instruct DocRes to perform various restoration tasks, we propose a novel visual prompt approach called Dynamic Task-Specific Prompt (DTSPrompt). The DTSPrompt for different tasks comprises distinct prior features, which are additional characteristics extracted from the input image. Beyond its role as a cue for task-specific execution, DTSPrompt can also serve as supplementary information to enhance the model's performance. Moreover, DTSPrompt is more flexible than prior visual prompt approaches as it can be seamlessly applied and adapted to inputs with high and variable resolutions. Experimental results demonstrate that DocRes achieves competitive or superior performance compared to existing state-of-the-art task-specific models. This underscores the potential of DocRes across a broader spectrum of document image restoration tasks. The source code is publicly available at https://github.com/ZZZHANG- jx/DocRes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie 等AAAI 2025 · 被引用 39 次
- ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content DisentanglementKahim Wong, Jicheng Zhou, Haiwei Wu, Yain-Whar Si 等ICCV 2025 · 被引用 3 次
- EpiAgent: An Agent-Centric System for Ancient Inscription RestorationShipeng Zhu, Ang Chen, Na Nie, Pengfei Fang 等CVPR 2026 · 被引用 2 次
- ForCenNet: Foreground-Centric Network for Document Image RectificationPeng Cai, Qiang Li, Kaicheng Yang, Dong Guo 等ICCV 2025 · 被引用 1 次
- Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionFangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Restormer: Efficient Transformer for High-Resolution Image RestorationSyed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat 等CVPR 2022 · 被引用 3,348 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- All-In-One Image Restoration for Unknown CorruptionBoyun Li, Xiao Liu, Peng Hu, Zhongqin Wu 等CVPR 2022 · 被引用 338 次
- A Unified Sequence Interface for Vision TasksTing Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin 等NeurIPS 2022 · 被引用 201 次
相关 Paper
- Learning Domain-Aware Task Prompt Representations for Multi-Domain All-in-One Image RestorationGuanglu Dong, Chunlei Li, Chao Ren, Jingliang Hu 等ICLR 2026 · 被引用 7 次
- PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionCong Wang, Jinshan Pan, Wei Wang, Jiangxin Dong 等NeurIPS 2023 · 被引用 109 次
- MMDIR: Multimodal Instruction-Driven Framework for Mixed-Degradation Document Image RestorationHeng Li, Xingyuan Wang, Yang Fan, Yunan Zhang 等CVPR 2026
- PromptIR: Prompting for All-in-One Image RestorationVaishnav Potlapalli, Syed Waqas Zamir, Salman H. Khan, Fahad Shahbaz KhanNeurIPS 2023 · 被引用 386 次
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather RemovalHanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang 等ACM MM 2025
