Visual Template Inference for Data Extraction from Documents
Yiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung, Aditya G. Parameswaran
摘要
Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, tax documents, financial reports, and purchase orders. Effective data extraction from these documents is crucial to support downstream analytical tasks. Current data extraction tools often struggle with complex document layouts, incur high latency and/or cost on large datasets, and require significant human effort. The key insight of our tool, TWIX, is to infer the underlying template used to create such documents, and then extract the data, rather than extracting directly from documents. To do so, TWIX first infers the underlying fields, such as columns of tabular portions or keys in co-located key-value pairs, by leveraging their consistent location patterns (e.g., two fields in the same template repeatedly co-occur within a fixed distance apart across multiple records). TWIX then assembles these fields into a template by enforcing visual constraints, such as vertically aligning table rows with their column headers for tabular regions, and horizontally aligning keys with their values for key-value pairs. TWIX then uses this inferred template to accurately and efficiently extract data from templatized documents at a low cost. On one benchmark with 34 diverse real-world datasets, TWIX outperforms state-of-the-art structured data extraction tools (Evaporate, Textract, and Azure Document Intelligence), and vision-based LLMs like GPT-4-Vision, by over 25% in precision and recall. Another benchmark with 30 large datasets demonstrates TWIX's scalability: it is 520X faster and 3,786X cheaper than the most competitive compared tool, for extracting data from large document collections with over 2000 pages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data LakesSimran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan 等VLDB 2024 · 被引用 165 次
- Towards End-to-End Unified Scene Text Detection and Layout AnalysisShangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco 等CVPR 2022 · 被引用 86 次
- DocETL: Agentic Query Rewriting and Evaluation for Complex Document ProcessingShreya Shankar, Tristan Chambers, Tarak Shah, Aditya G. Parameswaran 等VLDB 2025 · 被引用 62 次
- Form2Seq : A Framework for Higher-Order Form Structure ExtractionMilan Aggarwal, Hiresh Gupta, Mausoom Sarkar, Balaji KrishnamurthyEMNLP 2020 · 被引用 18 次
相关 Paper
- Glean: Structured Extractions from Templatic DocumentsSandeep Tata, Navneet Potti, James B. Wendt, Lauro Beltrão Costa 等VLDB 2021 · 被引用 17 次
- Querying Templatized Document Collections with Large Language ModelsYiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar 等ICDE 2025 · 被引用 4 次
- FieldSwap: Data Augmentation for Effective Form-Like Document ExtractionJing Xie, James B. Wendt, Yichao Zhou, Seth Ebner 等ICDE 2024
- LaTeX2Layout: High-Fidelity, Scalable Document Layout Annotation Pipeline for Layout DetectionFeijiang Han, Zelong Wang, Bowen Wang, Xinxin Liu 等AAAI 2026 · 被引用 4 次
- GistVis: Automatic Generation of Word-scale Visualizations from Data-rich DocumentsRuishi Zou, Yinqi Tang, Jingzhu Chen, Siyu Lu 等CHI 2025 · 被引用 8 次
