HRVDA: High-Resolution Visual Document Assistant
Chaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun, Linli Xu
Abstract
Leveraging vast training data, multimodal large language models (MLLMs) have demonstrated formidable general visual comprehension capabilities and achieved remarkable performance across various tasks. However, their performance in visual document understanding still leaves much room for improvement. This discrepancy is primarily attributed to the fact that visual document understanding is a fine-grained prediction task. In natural scenes, MLLMs typically use low-resolution images, leading to a substantial loss of visual information. Furthermore, general-purpose MLLMs do not excel in handling document-oriented instructions. In this paper, we propose a High-Resolution Visual Document Assistant (HRVDA), which bridges the gap between MLLMs and visual document understanding. This model employs a content filtering mechanism and an instruction filtering module to separately filter out the content-agnostic visual tokens and instruction-agnostic visual tokens, thereby achieving efficient model training and inference for high-resolution images. In addition, we construct a document-oriented visual instruction tuning dataset and apply a multi-stage training strategy to enhance the model's document modeling capabilities. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple document understanding datasets, while maintaining training efficiency and inference speed comparable to low-resolution models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4da8244d-5121-4bd4-80bc-e40d919c5372Cited by top-tier papers11
- DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual SlimmingJiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie et al.AAAI 2025 · 39 citations
- Recoverable Compression: A Multimodal Vision Token Recovery Mechanism Guided by Text InformationYi Chen, Jian Xu, Xu-Yao Zhang, Wen-Zhuo Liu et al.AAAI 2025 · 18 citations
- Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language ModelsYubo Wang, Chaohu Liu, Yanqiu Qu, Haoyu Cao et al.ACM MM 2024 · 10 citations
- AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference OptimizationChaohu Liu, Tianyi Gui, Yu Liu, Linli XuICLR 2026 · 9 citations
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of ExpertsYongXiang Hua, Haoyu Cao, Zhou Tao, Bocheng Li et al.ACM MM 2025 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Hierarchical Visual Feature Aggregation for OCR-Free Document UnderstandingJaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung HanNeurIPS 2024 · 19 citations
- InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with InstructionsRyota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito et al.AAAI 2024 · 39 citations
- LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningZebin You, Shen Nie, Xiaolu Zhang, JUN ZHOU et al.CVPR 2026 · 154 citations
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image UnderstandingFan Yang, Xingping Dong, Xin Yu, Wenhan Luo et al.CVPR 2026 · 1 citation
- LLaRA: Supercharging Robot Learning Data for Vision-Language PolicyXiang Li, Cristina Mata, Jongwoo Park, Kumara Kahatapitiya et al.ICLR 2025 · 2 citations
