Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Jaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung Han
Abstract
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1efaf2f5-9176-4af9-9cf4-f645e8f62c2dCited by top-tier papers2
- UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen et al.ACL 2026 · 3 citations
- Marten: Visual Question Answering with Mask Generation for Multi-modal Document UnderstandingZining Wang, Tongkun Guan, Pei Fu, Chen Duan et al.CVPR 2025
Builds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- HRVDA: High-Resolution Visual Document AssistantChaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang et al.CVPR 2024 · 10 citations
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingWenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang et al.CVPR 2025
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham et al.CVPR 2025
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng et al.CVPR 2024 · 39 citations
- SelfDoc: Self-Supervised Document Representation LearningPeizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu et al.CVPR 2021
