Form2Seq : A Framework for Higher-Order Form Structure Extraction
Milan Aggarwal, Hiresh Gupta, Mausoom Sarkar, Balaji Krishnamurthy
摘要
Document structure extraction has been a widely researched area for decades with recent works performing it as a semantic segmentation task over document images using fullyconvolution networks. Such methods are limited by image resolution due to which they fail to disambiguate structures in dense regions which appear commonly in forms. To mitigate this, we propose Form2Seq, a novel sequenceto-sequence (Seq2Seq) inspired framework for structure extraction using text, with a specific focus on forms, which leverages relative spatial arrangement of structures. We discuss two tasks; 1) Classification of low-level constituent elements (TextBlock and empty fillable Widget) into ten types such as field captions, list items, and others; 2) Grouping lower-level elements into higher-order constructs, such as Text Fields, ChoiceFields and ChoiceGroups, used as information collection mechanism in forms. To achieve this, we arrange the constituent elements linearly in natural reading order, feed their spatial and textual representations to Seq2Seq framework, which sequentially outputs prediction of each element depending on the final task. We modify Seq2Seq for grouping task and discuss improvements obtained through cascaded end-to-end training of two tasks versus training in isolation. Experimental results show the effectiveness of our text-based approach achieving an accuracy of 90% on classification task and an F1 of 75.82, 86.01, 61.63 on groups discussed above respectively, outperforming segmentation baselines. Further we show our framework achieves state of the results for table structure recognition on ICDAR 2013 dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot 等ACL 2022 · 被引用 90 次
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng 等WWW 2022 · 被引用 88 次
- MUSTIE: Multimodal Structural Transformer for Web Information ExtractionQifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng 等ACL 2023 · 被引用 16 次
- FormNetV2: Multimodal Graph Contrastive Learning for Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Hao Zhang, Timothy Dozat 等ACL 2023 · 被引用 7 次
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsYichao Zhou, James B. Wendt, Navneet Potti, Jing Xie 等EMNLP 2023
它引用的顶会 Paper2
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
相关 Paper
- Improving Table Structure Recognition with Visual-Alignment Sequential Coordinate ModelingYongshuai Huang, Ning Lu, Dapeng Chen, Yibo Li 等CVPR 2023
- StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingYuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 等ICLR 2023 · 被引用 18 次
- Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event ExtractionYaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han 等ACL 2021
- DocFormerv2: Local Features for Document UnderstandingSrikar Appalaraju, Peng Tang, Qi Dong, Nishant Sankaran 等AAAI 2024 · 被引用 68 次
- Seg2Act: Global Context-aware Action Generation for Document Logical StructuringZichao Li, Shaojie He, Meng Liao, Xuanang Chen 等EMNLP 2024 · 被引用 1 次
