Entity Relation Extraction as Dependency Parsing in Visually Rich Documents
Yue Zhang, Bo Zhang, Rui Wang, Junjie Cao, Chen Li, Zuyi Bao
Abstract
Previous works on key information extraction from visually rich documents (VRDs) mainly focus on labeling the text within each bounding box (i.e., semantic entity), while the relations in-between are largely unexplored. In this paper, we adapt the popular dependency parsing model, the biaffine parser, to this entity relation extraction task. Being different from the original dependency parsing model which recognizes dependency relations between words, we identify relations between groups of words with layout information instead. We have compared different representations of the semantic entity, different VRD encoders, and different relation decoders. For the model training, we explore multi-task learning to combine entity labeling and relation extraction tasks; and for the evaluation, we conduct experiments on different datasets with filtering and augmentation. The results demonstrate that our proposed model achieves 65.96% F1 score on the FUNSD dataset. As for the realworld application, our model has been applied to the in-house customs data, achieving reliable performance in the production setting. * Corresponding author. The author's contributions were carried out while at Alibaba Group. His current affiliation is Vipshop (China) Co., Ltd.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path PredictionChong Zhang, Ya Guo, Yi Tu, Huan Chen et al.EMNLP 2023 · 19 citations
- A Question-Answering Approach to Key Value Pair Extraction from Form-Like Document ImagesKai Hu, Zhuoyuan Wu, Zhuoyao Zhong, Weihong Lin et al.AAAI 2023 · 15 citations
- Modeling Layout Reading Order as Ordering Relations for Visually-rich Document UnderstandingChong Zhang, Yi Tu, Yixi Zhao, Chenshu Yuan et al.EMNLP 2024 · 4 citations
- GeoLayoutLM: Geometric Pre-training for Visual Information ExtractionChuwei Luo, Changxu Cheng, Qi Zheng, Cong YaoCVPR 2023
Builds on3
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt et al.ACL 2020 · 111 citations
- LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACL 2021
Related papers
- StrucTexT: Structured Text Understanding with Multi-Modal TransformersYulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin et al.ACM MM 2021 · 124 citations
- VRDSynth: Synthesizing Programs for Multilingual Visually Rich Document Information ExtractionThanh-Dat Nguyen, Tung Do-Viet, Hung Nguyen-Duy, Tuan-Hai Luu et al.ISSTA 2024 · 1 citation
- Relation-Rich Visual Document Generator for Visual Information ExtractionZi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu et al.CVPR 2025
- Relational Representation Learning in Visually-Rich DocumentsXin Li, Yan Zheng, Yiqing Hu, Haoyu Cao et al.ACM MM 2022 · 7 citations
- Multi-View Consistency for Relation Extraction via Mutual Information and Structure PredictionAmir Pouran Ben Veyseh, Franck Dernoncourt, My Tra Thai, Dejing Dou et al.AAAI 2020 · 17 citations
