Vision Grid Transformer for Document Layout Analysis
Cheng Da, Chuwei Luo, Qi Zheng, Cong Yao
Abstract
Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pretrained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a twostream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D 4 LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet (95.7%→96.2%), DocBank (79.6%→84.1%), and D 4 LA (67.7%→68.8%). The code and models as well as the D 4 LA dataset will be made publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 945a2114-d5ce-43a5-b7df-60776ec04f41Cited by top-tier papers11
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng et al.CVPR 2024 · 39 citations
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement LearningYuqi Liu, Tianyuan Qu, Zhisheng Zhong, Bohao Peng et al.ICLR 2026 · 15 citations
- SAIL: Sample-Centric In-Context Learning for Document Information ExtractionJinyu Zhang, Zhiyuan You, Jize Wang, Xinyi LeAAAI 2025 · 7 citations
- ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction DataYufan Shen, Chuwei Luo, Zhaoqing Zhu, Yang Chen et al.AAAI 2025 · 6 citations
- LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout DetectionFeijiang Han, Zelong Wang, Bowen Wang, Xinxin Liu et al.AAAI 2026 · 4 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
Related papers
- LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACL 2021
- M2Doc: A Multi-Modal Fusion Approach for Document Layout AnalysisNing Zhang, Hiuyi Cheng, Jiayu Chen, Zongyuan Jiang et al.AAAI 2024 · 16 citations
- Enhancing Visually-Rich Document Understanding via Layout Structure ModelingQiwei Li, Zuchao Li, Xiantao Cai, Bo Du et al.ACM MM 2023 · 9 citations
- LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document UnderstandingYi Tu, Ya Guo, Huan Chen, Jinyang TangACL 2023 · 21 citations
- M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisHiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang et al.CVPR 2023
