DocKylin: A Large Multimodal Model for Visual Document Understanding with Efficient Visual Slimming
Jiaxin Zhang, Wentao Yang, Songxuan Lai, Zecheng Xie, Lianwen Jin
摘要
Current multimodal large language models (MLLMs) face significant challenges in visual document understanding (VDU) tasks due to the high resolution, dense text, and complex layouts typical of document images. These characteristics demand a high level of detail perception ability from MLLMs. While increasing input resolution improves detail perception capability, it also leads to longer sequences of visual tokens, increasing computational costs and straining the models' ability to handle long contexts. To address these challenges, we introduce DocKylin, a document-centric MLLM that performs visual content slimming at both the pixel and token levels, thereby reducing token sequence length in VDU scenarios. We introduce an Adaptive Pixel Slimming (APS) preprocessing module to perform pixel-level slimming, increasing the proportion of informative pixels. Moreover, we propose a novel Dynamic Token Slimming (DTS) module to conduct token-level slimming, filtering essential tokens and removing others to adaptively create a more compact visual sequence. Experiments demonstrate DocKylin's promising performance across various VDU benchmarks and the effectiveness of each component.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai 等ICLR 2026 · 被引用 28 次
- A Token-Level Text Image Foundation Model for Document UnderstandingTongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo 等ICCV 2025 · 被引用 5 次
- Boosting Document Parsing Efficiency and Performance with Coarse-to-Fine Visual ProcessingCheng Cui, Ting Sun, Suyin Liang, Tingquan Gao 等CVPR 2026 · 被引用 3 次
- QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without RetrainingKyle R. Chickering, Bangzheng Li, Muhao ChenICLR 2026 · 被引用 2 次
- DocThinker: Explainable Multimodal Large Language Models with Rule-Based Reinforcement Learning for Document UnderstandingWenwen Yu, Zhibo Yang, Yuliang Liu, Xiang BaiICCV 2025 · 被引用 2 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
相关 Paper
- DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document UnderstandingWenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang 等CVPR 2025
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal UnderstandingYuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen 等CVPR 2026 · 被引用 2 次
- DocVLM: Make Your VLM an Efficient ReaderMor Shpigel Nacson, Aviad Aberdam, Roy Ganz, Elad Ben-Avraham 等CVPR 2025
- Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical FindingsQiong Wu, Wenhao Lin, Yiyi Zhou, Weihao Ye 等NeurIPS 2025 · 被引用 16 次
- A More Word-like Image Tokenization for MLLMsHyun Lee, Hyemin Jeong, Yejin Kim, Hyungwook Choi 等CVPR 2026 · 被引用 2 次
