FieldSwap: Data Augmentation for Effective Form-Like Document Extraction
Jing Xie, James B. Wendt, Yichao Zhou, Seth Ebner, Sandeep Tata
摘要
Extracting structured data from visually rich documents like invoices, receipts, financial statements, and tax forms is key to automating many business workflows. However, building extraction models in this domain often demands a large collection of high-quality training examples. To address this challenge, we introduce FieldSwap, a novel data augmentation technique specifically designed for such extraction problems. FieldSwap generates synthetic training examples by replacing key phrases indicative of one field with those corresponding to another. Our experiments on five diverse datasets demonstrate that incorporating FieldSwap-augmented data into the training process can enhance model performance by 1–11 F1 points, particularly when dealing with limited training data (10–100 documents). Additionally, we propose algorithms for automatically inferring key phrases from the training data. Our findings indicate that FieldSwap is effective regardless of whether key phrases are manually provided by human experts or inferred automatically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie 等ICCV 2021 · 被引用 392 次
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 被引用 114 次
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
相关 Paper
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsYichao Zhou, James B. Wendt, Navneet Potti, Jing Xie 等EMNLP 2023
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa 等VLDB 2021 · 被引用 17 次
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung 等SIGMOD 2026
- Sim VQA: Exploring Simulated Environments for Visual Question AnsweringPaola Cascante-Bonilla, Hui Wu, Letao Wang, Rogério Feris 等CVPR 2022 · 被引用 28 次
- Data Augmentation with Adversarial Training for Cross-Lingual NLIXin Dong, Yaxin Zhu, Zuohui Fu, Dongkuan Xu 等ACL 2021
