FieldSwap: Data Augmentation for Effective Form-Like Document Extraction
Jing Xie, James B. Wendt, Yichao Zhou, Seth Ebner, Sandeep Tata
Abstract
Extracting structured data from visually rich documents like invoices, receipts, financial statements, and tax forms is key to automating many business workflows. However, building extraction models in this domain often demands a large collection of high-quality training examples. To address this challenge, we introduce FieldSwap, a novel data augmentation technique specifically designed for such extraction problems. FieldSwap generates synthetic training examples by replacing key phrases indicative of one field with those corresponding to another. Our experiments on five diverse datasets demonstrate that incorporating FieldSwap-augmented data into the training process can enhance model performance by 1–11 F1 points, particularly when dealing with limited training data (10–100 documents). Additionally, we propose algorithms for automatically inferring key phrases from the training data. Our findings indicate that FieldSwap is effective regardless of whether key phrases are manually provided by human experts or inferred automatically.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1fdf1a-0c48-4cb9-93c9-18a993a6bfaeBuilds on12
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- DocFormer: End-to-End Transformer for Document UnderstandingSrikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie et al.ICCV 2021 · 392 citations
- Robustness to Spurious Correlations in Text Classification via Automatically Generated CounterfactualsZhao Wang, Aron CulottaAAAI 2021 · 114 citations
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt et al.ACL 2020 · 111 citations
Related papers
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsYichao Zhou, James B. Wendt, Navneet Potti, Jing Xie et al.EMNLP 2023
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa et al.VLDB 2021 · 17 citations
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung et al.SIGMOD 2026
- Sim VQA: Exploring Simulated Environments for Visual Question AnsweringPaola Cascante-Bonilla, Hui Wu, Letao Wang, Rogério Feris et al.CVPR 2022 · 28 citations
- Data Augmentation with Adversarial Training for Cross-Lingual NLIXin Dong, Yaxin Zhu, Zuohui Fu, Dongkuan Xu et al.ACL 2021
