Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction
Chong Zhang, Ya Guo, Yi Tu, Huan Chen, Jinyang Tang, Huijia Zhu, Qi Zhang, Tao Gui
摘要
Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP. However, BIO-tagging scheme relies on the correct order of model inputs, which is not guaranteed in real-world NER on scanned VrDs where text are recognized and arranged by OCR systems. Such reading order issue hinders the accurate marking of entities by BIO-tagging scheme, making it impossible for sequence-labeling methods to predict correct named entities. To address the reading order issue, we introduce Token Path Prediction (TPP), a simple prediction head to predict entity mentions as token sequences within documents. Alternative to token classification, TPP models the document layout as a complete directed graph of tokens, and predicts token paths within the graph as entities. For better evaluation of VrD-NER systems, we also propose two revised benchmark datasets of NER on scanned documents which can reflect real-world scenarios. Experiment results demonstrate the effectiveness of our method, and suggest its potential to be a universal solution to various information extraction tasks on documents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu 等CVPR 2024 · 被引用 29 次
- PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair ExtractionZening Lin, Jiapeng Wang, Teng Li, Wenhui Liao 等ACM MM 2024 · 被引用 6 次
- UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual DocumentsYifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen 等ACL 2026 · 被引用 3 次
- DREAM: Document Reconstruction via End-to-end Autoregressive ModelXin Li, Mingming Gong, Yunfei Wu, Jianxin Dai 等ACM MM 2025 · 被引用 2 次
- UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsYi Tu, Chong Zhang, Ya Guo, Huan Chen 等ACM MM 2024 · 被引用 2 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Unified Named Entity Recognition as Word-Word Relation ClassificationJingye Li, Hao Fei, Jiang Liu, Shengqiong Wu 等AAAI 2022 · 被引用 340 次
- LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document UnderstandingJiapeng Wang, Lianwen Jin, Kai DingACL 2022 · 被引用 188 次
相关 Paper
- LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document UnderstandingYi Tu, Ya Guo, Huan Chen, Jinyang TangACL 2023 · 被引用 21 次
- Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan 等EMNLP 2024 · 被引用 4 次
- XYLayoutLM: Towards Layout-Aware Multimodal Networks For Visually-Rich Document UnderstandingZhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan 等CVPR 2022 · 被引用 82 次
- A Simple yet Effective Layout Token in Large Language Models for Document UnderstandingZhaoqing Zhu, Chuwei Luo, Zirui Shao, Feiyu Gao 等CVPR 2025
- MarkupLM: Pre-training of Text and Markup Language for Visually Rich Document UnderstandingJunlong Li, Yiheng Xu, Lei Cui, Furu WeiACL 2022 · 被引用 75 次
