Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction Models
Yichao Zhou, James B. Wendt, Navneet Potti, Jing Xie, Sandeep Tata
Abstract
Building automatic extraction models for visually rich documents like invoices, receipts, bills, tax forms, etc. has received significant attention lately. A key bottleneck in developing extraction models for new document types is the cost of acquiring the several thousand high-quality labeled documents that are needed to train a model with acceptable accuracy. In this paper, we propose selective labeling as a solution to this problem. The key insight is to simplify the labeling task to provide “yes/no” labels for candidate extractions predicted by a model trained on partially labeled documents. We combine this with a custom active learning strategy to find the predictions that the model is most uncertain about. We show through experiments on document types drawn from 3 different domains that selective labeling can reduce the cost of acquiring labeled data by 10× with a negligible loss in accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 128 citations
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt et al.ACL 2020 · 111 citations
Related papers
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao et al.AAAI 2020 · 27 citations
- FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionJing Xie, James B. Wendt, Yichao Zhou, Seth Ebner et al.ICDE 2024
- SEL-BALD: Deep Bayesian Active Learning with Selective LabelsRuijiang Gao, Mingzhang Yin, Maytal Saar-TsechanskyNeurIPS 2024 · 4 citations
- Adapting Coreference Resolution Models through Active LearningMichelle Yuan, Patrick Xia, Chandler May, Benjamin Van Durme et al.ACL 2022 · 20 citations
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa et al.VLDB 2021 · 17 citations
