One-shot Text Field labeling using Attention and Belief Propagation for Structure Information Extraction
Mengli Cheng, Minghui Qiu, Xing Shi, Jun Huang, Wei Lin
摘要
Structured information extraction from document images usually consists of three steps: text detection, text recognition, and text field labeling. While text detection and text recognition have been heavily studied and improved a lot in literature, text field labeling is less explored and still faces many challenges. Existing learning based methods for text labeling task usually require a large amount of labeled examples to train a specific model for each type of document. However, collecting large amounts of document images and labeling them is difficult and sometimes impossible due to privacy issues. Deploying separate models for each type of document also consumes a lot of resources. Facing these challenges, we explore one-shot learning for the text field labeling task. Existing one-shot learning methods for the task are mostly rule-based and have difficulty in labeling fields in crowded regions with few landmarks and fields consisting of multiple separate text regions. To alleviate these problems, we proposed a novel deep end-to-end trainable approach for one-shot text field labeling, which makes use of attention mechanism to transfer the layout information between document images. We further applied conditional random field on the transferred layout information for the refinement of field labeling. We collected and annotated a real-world one-shot field labeling dataset with a large variety of document types and conducted extensive experiments to examine the effectiveness of the proposed model. To stimulate research in this direction, the collected dataset and the one-shot model will be released (https://github.com/AlibabaPAI/one_shot_text_labeling).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- StrucTexT: Structured Text Understanding with Multi-Modal TransformersYulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin 等ACM MM 2021 · 被引用 124 次
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
- Query-driven Generative Network for Document Information Extraction in the WildHaoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang 等ACM MM 2022 · 被引用 13 次
- Enhancing Visually-Rich Document Understanding via Layout Structure ModelingQiwei Li, Zuchao Li, Xiantao Cai, Bo Du 等ACM MM 2023 · 被引用 9 次
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer 等PLDI 2022 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- Pyramid Graph Networks With Connection Attentions for Region-Based One-Shot Semantic SegmentationChi Zhang, Guosheng Lin, Fayao Liu, Jiushuang Guo 等ICCV 2019 · 被引用 351 次
- LT-Net: Label Transfer by Learning Reversible Voxel-Wise Correspondence for One-Shot Medical Image SegmentationShuxin Wang, Shilei Cao, Dong Wei, Renzhen Wang 等CVPR 2020
- End-to-End Modeling via Information Tree for One-Shot Natural Language Spatial Video GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang 等ACL 2022 · 被引用 46 次
- Adaptive Image Transformer for One-Shot Object DetectionDing-Jie Chen, He-Yen Hsieh, Tyng-Luh LiuCVPR 2021
- Few-shot Slot Tagging with Collapsed Dependency Transfer and Label-enhanced Task-adaptive Projection NetworkYutai Hou, Wanxiang Che, Yongkui Lai, Zhihan Zhou 等ACL 2020 · 被引用 186 次
