LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding
Yi Tu, Ya Guo, Huan Chen, Jinyang Tang
摘要
Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant performance gains in this field. The major challenge is how to fusion the different modalities (text, layout, and image) of the documents in a unified model with different pre-training tasks. This paper focuses on improving text-layout interactions and proposes a novel multi-modal pre-training model, LayoutMask. LayoutMask uses local 1D position, instead of global 1D position, as layout input and has two pre-training objectives: (1) Masked Language Modeling: predicting masked tokens with two novel masking strategies; (2) Masked Position Modeling: predicting masked 2D positions to improve layout representation learning. LayoutMask can enhance the interactions between text and layout modalities in a unified model and produce adaptive and robust multimodal representations for downstream tasks. Experimental results show that our proposed method can achieve state-of-the-art results on a wide variety of VrDU problems, including form understanding, receipt understanding, and document image classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LayoutLLM: Layout Instruction Tuning with Large Language Models for Document UnderstandingChuwei Luo, Yufan Shen, Zhaoqing Zhu, Qi Zheng 等CVPR 2024 · 被引用 39 次
- Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path PredictionChong Zhang, Ya Guo, Yi Tu, Huan Chen 等EMNLP 2023 · 被引用 19 次
- DocMamba: Efficient Document Pre-training with State Space ModelPengfei Hu, Zhenrong Zhang, Jiefeng Ma, Shuhang Liu 等AAAI 2025 · 被引用 4 次
- Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan 等EMNLP 2024 · 被引用 4 次
- UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsYi Tu, Chong Zhang, Ya Guo, Huan Chen 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper12
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
相关 Paper
- LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui 等ACL 2021
- XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingZhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan 等CVPR 2022 · 被引用 82 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
- StructuralLM: Structural Pre-training for Form UnderstandingChenliang Li, Bin Bi, Ming Yan, Wei Wang 等ACL 2021
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 被引用 75 次
