One-shot Text Field labeling using Attention and Belief Propagation for Structure Information Extraction
Mengli Cheng, Minghui Qiu, Xing Shi, Jun Huang, Wei Lin
Abstract
Structured information extraction from document images usually consists of three steps: text detection, text recognition, and text field labeling. While text detection and text recognition have been heavily studied and improved a lot in literature, text field labeling is less explored and still faces many challenges. Existing learning based methods for text labeling task usually require a large amount of labeled examples to train a specific model for each type of document. However, collecting large amounts of document images and labeling them is difficult and sometimes impossible due to privacy issues. Deploying separate models for each type of document also consumes a lot of resources. Facing these challenges, we explore one-shot learning for the text field labeling task. Existing one-shot learning methods for the task are mostly rule-based and have difficulty in labeling fields in crowded regions with few landmarks and fields consisting of multiple separate text regions. To alleviate these problems, we proposed a novel deep end-to-end trainable approach for one-shot text field labeling, which makes use of attention mechanism to transfer the layout information between document images. We further applied conditional random field on the transferred layout information for the refinement of field labeling. We collected and annotated a real-world one-shot field labeling dataset with a large variety of document types and conducted extensive experiments to examine the effectiveness of the proposed model. To stimulate research in this direction, the collected dataset and the one-shot model will be released (https://github.com/AlibabaPAI/one_shot_text_labeling).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- StrucTexT: Structured Text Understanding with Multi-Modal TransformersYulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin et al.ACM MM 2021 · 124 citations
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
- Query-driven Generative Network for Document Information Extraction in the WildHaoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang et al.ACM MM 2022 · 13 citations
- Enhancing Visually-Rich Document Understanding via Layout Structure ModelingQiwei Li, Zuchao Li, Xiantao Cai, Bo Du et al.ACM MM 2023 · 9 citations
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer et al.PLDI 2022 · 2 citations
Builds on1
Related papers
- Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic SegmentationChi Zhang, Guosheng Lin, Fayao Liu, Jiushuang Guo et al.ICCV 2019 · 351 citations
- LT-Net: Label Transfer by Learning Reversible Voxel-Wise Correspondence for One-Shot Medical Image SegmentationShuxin Wang, Shilei Cao, Dong Wei, Renzhen Wang et al.CVPR 2020
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACL 2022 · 46 citations
- Adaptive Image Transformer for One-Shot Object DetectionDing-Jie Chen, He-Yen Hsieh, Tyng-Luh LiuCVPR 2021
- Few-shot Slot Tagging with Collapsed Dependency Transfer and Label-enhanced Task-adaptive Projection NetworkYutai Hou, Wanxiang Che, Yongkui Lai, Zhihan Zhou et al.ACL 2020 · 186 citations
