Vision Grid Transformer for Document Layout Analysis
Cheng Da, Chuwei Luo, Qi Zheng, Cong Yao
摘要
Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those pretrained in a multi-modal fashion, usually rely on either textual features or visual features. Grid-based models for DLA are multi-modality but largely neglect the effect of pre-training. To fully leverage multi-modal information and exploit pre-training techniques to learn better representation for DLA, in this paper, we present VGT, a twostream Vision Grid Transformer, in which Grid Transformer (GiT) is proposed and pre-trained for 2D token-level and segment-level semantic understanding. Furthermore, a new dataset named D 4 LA, which is so far the most diverse and detailed manually-annotated benchmark for document layout analysis, is curated and released. Experiment results have illustrated that the proposed VGT model achieves new state-of-the-art results on DLA tasks, e.g. PubLayNet (95.7%→96.2%), DocBank (79.6%→84.1%), and D 4 LA (67.7%→68.8%). The code and models as well as the D 4 LA dataset will be made publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 等CVPR 2024 · 被引用 39 次
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement LearningYuqi Liu, Tianyuan Qu, Zhisheng Zhong, Bohao Peng 等ICLR 2026 · 被引用 15 次
- SAIL: Sample-Centric In-Context Learning for Document Information ExtractionJinyu Zhang, Zhiyuan You, Jize Wang, Xinyi LeAAAI 2025 · 被引用 7 次
- ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction DataYufan Shen, Chuwei Luo, Zhaoqing Zhu, Yang Chen 等AAAI 2025 · 被引用 6 次
- LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout DetectionFeijiang Han, Zelong Wang, Bowen Wang, Xinxin Liu 等AAAI 2026 · 被引用 4 次
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li 等CVPR 2022 · 被引用 481 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
相关 Paper
- LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui 等ACL 2021
- M2Doc: A Multi-Modal Fusion Approach for Document Layout AnalysisNing Zhang, Hiuyi Cheng, Jiayu Chen, Zongyuan Jiang 等AAAI 2024 · 被引用 16 次
- Enhancing Visually-Rich Document Understanding via Layout Structure ModelingQiwei Li, Zuchao Li, Xiantao Cai, Bo Du 等ACM MM 2023 · 被引用 9 次
- LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document UnderstandingYi Tu, Ya Guo, Huan Chen, Jinyang TangACL 2023 · 被引用 21 次
- M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout AnalysisHiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang 等CVPR 2023
