Glean: Structured Extractions from Templatic Documents
Sandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa, Marc Najork, Beliz Gunel
Abstract
Extracting structured information from templatic documents is an important problem with the potential to automate many real-world business workflows such as payment, procurement, and payroll. The core challenge is that such documents can be laid out in virtually infinitely different ways. A good solution to this problem is one that generalizes well not only to known templates such as invoices from a known vendor, but also to unseen ones.
We developed a system called Glean to tackle this problem. Given a target schema for a document type and some labeled documents of that type, Glean uses machine learning to automatically extract structured information from other documents of that type. In this paper, we describe the overall architecture of Glean, and discuss three key data management challenges : 1) managing the quality of ground truth data, 2) generating training data for the machine learning model using labeled documents, and 3) building tools that help a developer rapidly build and improve a model for a given document type. Through empirical studies on a real-world dataset, we show that these data management techniques allow us to train a model that is over 5 F1 points better than the exact same model architecture without the techniques we describe. We argue that for such information-extraction problems, designing abstractions that carefully manage the training data is at least as important as choosing a good model architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdec2909-5dc1-4114-a1d6-0d370c702712Cited by top-tier papers4
- Learning Structural Co-occurrences for Structured Web Data Extraction in Low-Resource SettingsZhenyu Zhang, Bowen Yu, Tingwen Liu, Tianyun Liu et al.WWW 2023 · 7 citations
- Selective Labeling: How to Radically Lower Data-Labeling Costs for Document Extraction ModelsYichao Zhou, James B. Wendt, Navneet Potti, Jing Xie et al.EMNLP 2023
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung et al.SIGMOD 2026
- FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionJing Xie, James B. Wendt, Yichao Zhou, Seth Ebner et al.ICDE 2024
Builds on2
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt et al.ACL 2020 · 111 citations
Related papers
- Landmarks and regions: a robust approach to data extractionSuresh Parthasarathy, Lincy Pattanaik, Anirudh Khatry, Arun Iyer et al.PLDI 2022 · 2 citations
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan et al.VLDB 2024 · 165 citations
- Querying Templatized Document Collections with Large Language ModelsYiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar et al.ICDE 2025 · 4 citations
- Query-driven Generative Network for Document Information Extraction in the WildHaoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang et al.ACM MM 2022 · 13 citations
- From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document UnderstandingYandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo et al.ACL 2026
