Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction Models
Yichao Zhou, James B. Wendt, Navneet Potti, Jing Xie, Sandeep Tata
摘要
Building automatic extraction models for visually rich documents like invoices, receipts, bills, tax forms, etc. has received significant attention lately. A key bottleneck in developing extraction models for new document types is the cost of acquiring the several thousand high-quality labeled documents that are needed to train a model with acceptable accuracy. In this paper, we propose selective labeling as a solution to this problem. The key insight is to simplify the labeling task to provide “yes/no” labels for candidate extractions predicted by a model trained on partially labeled documents. We combine this with a custom active learning strategy to find the predictions that the model is most uncertain about. We show through experiments on document types drawn from 3 different domains that selective labeling can reduce the cost of acquiring labeled data by 10× with a negligible loss in accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
相关 Paper
- Active Learning with Query Generation for Cost-Effective Text ClassificationYifan Yan, Sheng-Jun Huang, Shaoyi Chen, Meng Liao 等AAAI 2020 · 被引用 27 次
- FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionJing Xie, James B. Wendt, Yichao Zhou, Seth Ebner 等ICDE 2024
- SEL-BALD: Deep Bayesian Active Learning with Selective LabelsRuijiang Gao, Mingzhang Yin, Maytal Saar-TsechanskyNeurIPS 2024 · 被引用 4 次
- Adapting Coreference Resolution Models through Active LearningMichelle Yuan, Patrick Xia, Chandler May, Benjamin Van Durme 等ACL 2022 · 被引用 20 次
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa 等VLDB 2021 · 被引用 17 次
