UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
Yifan Ji, Zhipeng Xu, Zhenghao Liu, Zulong Chen, Qian Zhang, Zhibo Yang, Junyang Lin, Yu Gu, Ge Yu, Maosong Sun
摘要
Key Information Extraction (KIE) from realworld documents remains challenging due to substantial variations in layout structures, visual quality, and task-specific information requirements. Recent Large Multimodal Models (LMMs) have shown promising potential for performing end-to-end KIE directly from document images. To enable a comprehensive and systematic evaluation across realistic and diverse application scenarios, we introduce UNIKIE-BENCH, a unified benchmark designed to rigorously evaluate the KIE capabilities of LMMs. UNIKIE-BENCH consists of two complementary tracks: a constrainedcategory KIE track with scenario-predefined schemas that reflect practical application needs, and an open-category KIE track that extracts any key information that is explicitly present in the document. Experiments on 15 stateof-the-art LMMs reveal substantial performance degradation under diverse schema definitions, long-tail key fields, and complex layouts, along with pronounced performance disparities across different document types and scenarios. These findings underscore persistent challenges in grounding accuracy and layout-aware reasoning for LMM-based KIE. All codes and datasets are available at https: //github.com/NEUIR/UNIKIE-BENCH . * indicates equal contribution. † indicates corresponding author.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 等ACM MM 2022 · 被引用 606 次
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang 等KDD 2020 · 被引用 575 次
- Towards Robust Visual Information Extraction in Real World: New Dataset and Novel SolutionJiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang 等AAAI 2021 · 被引用 97 次
- FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information ExtractionChen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot 等ACL 2022 · 被引用 90 次
- DocLLM: A Layout-Aware Generative Language Model for Multimodal Document UnderstandingDongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma 等ACL 2024 · 被引用 37 次
相关 Paper
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in LiteracyZhibo Yang, Jun Tang, Zhaohai Li, Pengfei Wang 等ICCV 2025 · 被引用 15 次
- UNICBench: UNIfied Counting Benchmark for MLLMChenggang Rong, Tao Han, Zhiyuan Zhao, Yaowu Fan 等CVPR 2026 · 被引用 3 次
- OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive AnnotationsLinke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu 等CVPR 2025
- MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual KnowledgeYuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi 等ICLR 2025
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu 等ICLR 2026 · 被引用 15 次
