Improving Information Extraction from Visually Rich Documents using Visual Span Representations
Ritesh Sarkhel, Arnab Nandi
摘要
Along with textual content, visual features play an essential role in the semantics of visually rich documents. Information extraction (IE) tasks perform poorly on these documents if these visual cues are not taken into account. In this paper, we present Artemis - a visually aware, machine-learning-based IE method for heterogeneous visually rich documents. Artemis represents a visual span in a document by jointly encoding its visual and textual context for IE tasks. Our main contribution is two-fold. First, we develop a deep-learning model that identifies the local context boundary of a visual span with minimal human-labeling. Second, we describe a deep neural network that encodes the multimodal context of a visual span into a fixed-length vector by taking its textual and layout-specific features into account. It identifies the visual span(s) containing a named entity by leveraging this learned representation followed by an inference task. We evaluate Artemis on four heterogeneous datasets from different domains over a suite of information extraction tasks. Results show that it outperforms state-of-the-art text-based methods by up to 17 points in F1-score.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-PagesRitesh Sarkhel, Binxuan Huang, Colin Lockard, Prashant ShiralkarVLDB 2023 · 被引用 11 次
- Visual Template Inference for Data Extraction from DocumentsYiming Lin, Mawil Hasan, Rohan Kosalge, Alvin Cheung 等SIGMOD 2026
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
它引用的顶会 Paper1
相关 Paper
- UNER: A Unified Prediction Head for Named Entity Recognition in Visually-rich DocumentsYi Tu, Chong Zhang, Ya Guo, Huan Chen 等ACM MM 2024 · 被引用 2 次
- Hypergraph based Understanding for Document Semantic Entity RecognitionQiwei Li, Zuchao Li, Ping Wang, Haojun Ai 等ACL 2024 · 被引用 2 次
- Improving Multimodal Named Entity Recognition via Entity Span Detection with Unified Multimodal TransformerJianfei Yu, Jing Jiang, Li Yang, Rui XiaACL 2020 · 被引用 260 次
- Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual SegmentsAniket Bhattacharyya, Anurag Tripathi, Ujjal Das, Archan Karmakar 等ACL 2025
- DocTr: Document Transformer for Structured Information Extraction in DocumentsHaofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal 等ICCV 2023 · 被引用 26 次
