Uni-DocRobust: Universal Plug-and-Play Robustness Enhancement for Multi-modal LLMs via Feature Restoration
Yuxuan Zhou, Baole Wei, Xingjian Hu, Haowei Chen, Yu Li, Xingyue Lin, Liangcai Gao, Zhi Tang
摘要
Real-world degradations, such as noise, blur, and low resolution, significantly impair the performance of Multi-modal Large Language Models (MLLMs) in document understanding tasks. Despite recent advancements, progress in this field remains stifled by two critical bottlenecks: the scarcity of large-scale, aligned training data necessary for learning robustness, and the lack of transferable restoration solutions across diverse MLLM architectures. To bridge the data gap, we first present DocRobust-VQA, a large-scale dataset explicitly constructed to support robustness training. Comprising 189K aligned clean/corrupted document image pairs and 417K QA pairs, it provides the first substantial corpus for fine-tuning MLLMs to handle varying degradation conditions. Leveraging this data, we propose Uni-DocRobust, a universal plug-and-play framework that decouples restoration capabilities from specific visual encoders. Our method employs a frozen Universal Restoration Core pre-trained in a canonical feature space via multi-teacher distillation, which can be seamlessly integrated into target MLLMs (e.g., Qwen-VL, InternVL) through lightweight Feature Adapters. Extensive experiments demonstrate that Uni-DocRobust significantly enhances robust performance on MLLMs and enables a cost-effective ``pre-train once, deploy everywhere'' paradigm for robust MLLM deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 等NeurIPS 2023 · 被引用 404 次
相关 Paper
- Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?Jiaqi Tang, Jianmin Chen, Youyang Zhai, Wei Wei 等ICML 2026 · 被引用 1 次
- Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich DocumentsDavide Napolitano, Luca Cagliero, Fabrizio BattiloroAAAI 2026
- Robust-R1: Degradation-Aware Reasoning for Robust Visual UnderstandingJiaqi Tang, Jianmin Chen, Wei Wei, Xiaogang Xu 等AAAI 2026 · 被引用 4 次
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham 等CVPR 2025
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language ModelsZhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen 等NeurIPS 2025 · 被引用 14 次
