Glean: Structured Extractions from Templatic Documents
Sandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa, Marc Najork, Beliz Gunel
摘要
Extracting structured information from templatic documents is an important problem with the potential to automate many real-world business workflows such as payment, procurement, and payroll. The core challenge is that such documents can be laid out in virtually infinitely different ways. A good solution to this problem is one that generalizes well not only to known templates such as invoices from a known vendor, but also to unseen ones.
We developed a system called Glean to tackle this problem. Given a target schema for a document type and some labeled documents of that type, Glean uses machine learning to automatically extract structured information from other documents of that type. In this paper, we describe the overall architecture of Glean, and discuss three key data management challenges : 1) managing the quality of ground truth data, 2) generating training data for the machine learning model using labeled documents, and 3) building tools that help a developer rapidly build and improve a model for a given document type. Through empirical studies on a real-world dataset, we show that these data management techniques allow us to train a model that is over 5 F1 points better than the exact same model architecture without the techniques we describe. We argue that for such information-extraction problems, designing abstractions that carefully manage the training data is at least as important as choosing a good model architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu 等WWW 2023 · 被引用 7 次
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsYichao Zhou, James B. Wendt, Navneet Potti, Jing Xie 等EMNLP 2023
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung 等SIGMOD 2026
- FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionJing Xie, James B. Wendt, Yichao Zhou, Seth Ebner 等ICDE 2024
它引用的顶会 Paper2
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
相关 Paper
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer 等PLDI 2022 · 被引用 2 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- Querying Templatized Document Collections with Large Language ModelsYiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar 等ICDE 2025 · 被引用 4 次
- Query-driven Generative Network for Document Information Extraction in the WildHaoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang 等ACM MM 2022 · 被引用 13 次
- From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document UnderstandingYandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo 等ACL 2026
