Query-driven Generative Network for Document Information Extraction in the Wild
Haoyu Cao, Xin Li, Jiefeng Ma, Deqiang Jiang, Antai Guo, Yiqing Hu, Hao Liu, Yinsong Liu, Bo Ren
摘要
This paper focuses on solving Document Information Extraction (DIE) in the wild problem, which is rarely explored before. In contrast to existing studies mainly tailored for document cases in known templates with predefined layouts and keys under the ideal input without OCR errors involved, we aim to build up a more practical DIE paradigm for real-world scenarios where input document images may contain unknown layouts and keys in the scenes of the problematic OCR results. To achieve this goal, we propose a novel architecture, termed Query-driven Generative Network (QGN), which is equipped with two consecutive modules, i.e., Layout Context-aware Module (LCM) and Structured Generation Module (SGM). Given a document image with unseen layouts and fields, the former LCM yields the value prefix candidates serving as the query prompts for the SGM to generate the final key-value pairs even with OCR noise. To further investigate the potential of our method, we create a new large-scale dataset, named LArge-scale STructured Documents (LastDoc4000), containing 4,000 documents with 1,511 layouts and 3,500 different keys. In experiments, we demonstrate that our QGN consistently achieves the best F1-score on the new LastDoc4000 dataset by at most 30.32% absolute improvement. A more comprehensive experimental analysis and experiments on other public benchmarks also verify the effectiveness and robustness of our proposed method for the wild DIE task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information ExtractionJiabang He, Lei Wang, Yi Hu, Ning Liu 等ICCV 2023 · 被引用 61 次
- OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table RecognitionJianqiang Wan, Sibo Song, Wenwen Yu, Yuliang Liu 等CVPR 2024 · 被引用 29 次
- Attention Where It Matters: Rethinking Visual Document Understanding with Selective Region ConcentrationHaoyu Cao, Changcun Bao, Chaohu Liu, Huang Chen 等ICCV 2023 · 被引用 21 次
- HRVDA: High-Resolution Visual Document AssistantChaohu Liu, Kun Yin, Haoyu Cao, Xinghua Jiang 等CVPR 2024 · 被引用 10 次
- PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair ExtractionZening Lin, Jiapeng Wang, Teng Li, Wenhui Liao 等ACM MM 2024 · 被引用 6 次
它引用的顶会 Paper11
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from DocumentsTeakgyu Hong, Donghyun Kim, Mingi Ji, Wonseok Hwang 等AAAI 2022 · 被引用 186 次
- StrucTexT: Structured Text Understanding with Multi-Modal TransformersYulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin 等ACM MM 2021 · 被引用 124 次
- Representation Learning for Information Extraction from Form-like DocumentsBodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt 等ACL 2020 · 被引用 111 次
相关 Paper
- RDLNet: A Novel and Accurate Real-world Document Localization MethodYaqiang Wu, Zhen Xu, Yong Duan, Yanlai Wu 等ACM MM 2024 · 被引用 6 次
- Relation-Rich Visual Document Generator for Visual Information ExtractionZi-Han Jiang, Chien-Wei Lin, Wei-Hua Li, Hsuan-Tung Liu 等CVPR 2025
- Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual SegmentsAniket Bhattacharyya, Anurag Tripathi, Ujjal Das, Archan Karmakar 等ACL 2025
- UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsYi Tu, Chong Zhang, Ya Guo, Huan Chen 等ACM MM 2024 · 被引用 2 次
- SAIL: Sample-Centric In-Context Learning for Document Information ExtractionJinyu Zhang, Zhiyuan You, Jize Wang, Xinyi LeAAAI 2025 · 被引用 7 次
