Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Jaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung Han
摘要
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen 等ACL 2026 · 被引用 3 次
- Marten: Visual Question Answering with Mask Generation for Multi-modal Document UnderstandingZining Wang, Tongkun Guan, Pei Fu, Chen Duan 等CVPR 2025
它引用的顶会 Paper29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- HRVDA: High-Resolution Visual Document AssistantChaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang 等CVPR 2024 · 被引用 10 次
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingWenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang 等CVPR 2025
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham 等CVPR 2025
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 等CVPR 2024 · 被引用 39 次
- SelfDoc: Self-Supervised Document Representation LearningPeizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu 等CVPR 2021
