MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering
Teng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang, Chunlin Liu, Kaishun Wu, Nan Tang
摘要
Cross-Document Multi-entity Question Answering (MEQA) demands the integration of scattered information across documents to resolve complex queries involving entities, relationships, and contextual dependencies. Although Large Language Models (LLMs) and Retrieval-augmented Generation (RAG) systems show promise, their performance on crossdocument MEQA remains underexplored due to the absence of tailored benchmarks. To address this gap, we introduce MEBench, a scalable multi-document, multi-entity benchmark designed to systematically evaluate LLMs' capacity to retrieve, consolidate, and reason over scattered and dense information. Our benchmark comprises 4,780 questions which are systematically categorized into three primary categories: Comparative Reasoning, Statistical Reasoning and Relational Reasoning, further divided into eight distinct types, ensuring broad coverage of real-world multi-entity reasoning scenarios. Our experiments on state-of-theart LLMs reveal critical limitations: even advanced models achieve only 59% accuracy on MEBench. Our benchmark emphasizes the importance of completeness and factual precision of information extraction in MEQA tasks, using Entity-Attributed F1 (EA-F1) metric for granular evaluation of entity-level correctness and attribution validity. MEBench not only highlights systemic weaknesses in current LLM frameworks but also provides a foundation for advancing robust, entity-aware QA architectures. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah 等ISCA 2024 · 被引用 282 次
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi 等SIGIR 2024 · 被引用 182 次
- Natural Language to Visualization by Neural Machine TranslationYuyu Luo, Nan Tang, Guoliang Li, Jiawei Tang 等IEEE VIS 2021 · 被引用 145 次
- Synthesizing Natural Language to Visualization (NL2VIS) Benchmarks from NL2SQL BenchmarksYuyu Luo, Nan Tang, Guoliang Li, Chengliang Chai 等SIGMOD 2021 · 被引用 90 次
- Cost-Effective In-Context Learning for Entity Resolution: A Design Space ExplorationMeihao Fan, Xiaoyue Han, Ju Fan, Chengliang Chai 等ICDE 2024 · 被引用 40 次
相关 Paper
- Are We on the Right Way to Assess Document Retrieval-Augmented Generation?Wenxuan Shen, Mingjia Wang, Yaochen Wang, Dongping Chen 等AAAI 2026
- PRGB Benchmark: A Robust Placeholder-Assisted Algorithm for Benchmarking Retrieval-Augmented GenerationZhehao Tan, Yihan Jiao, Dan Yang, Junwei Liu 等AAAI 2026
- RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity CorporaHanjun Cho, Jay-Yoon LeeACL 2026 · 被引用 1 次
- M³GQA: A Multi-Entity Multi-Hop Multi-Setting Graph Question Answering BenchmarkBoci Peng, Yongchao Liu, Xiaohe Bo, Jiaxin Guo 等ACL 2025 · 被引用 1 次
- Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual DataSeiji Maekawa, Hayate Iso, Nikita BhutaniICLR 2025
