Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data
Seiji Maekawa, Hayate Iso, Nikita Bhutani
摘要
The rapid increase in textual information means we need more efficient methods to sift through, organize, and understand it all. While retrieval-augmented generation (RAG) models excel in accessing information from large document collections, they struggle with complex tasks that require aggregation and reasoning over information spanning across multiple documents-what we call holistic reasoning. Long-context language models (LCLMs) have great potential for managing large-scale documents, but their holistic reasoning capabilities remain unclear. In this work, we introduce HoloBench, a novel framework a novel framework that brings database reasoning operations into text-based contexts, making it easier to systematically evaluate how LCLMs handle holistic reasoning across large documents. Our approach adjusts key factors such as context length, information density, distribution of information, and query complexity to evaluate LCLMs comprehensively. Our experiments show that the amount of information in the context has a bigger influence on LCLM performance than the actual context length. Furthermore, the complexity of queries affects performance more than the amount of information, particularly for different types of queries. Interestingly, queries that involve finding maximum or minimum values are easier for LCLMs and are less affected by context length, even though they pose challenges for RAG systems. However, tasks requiring the aggregation of multiple pieces of information show a noticeable drop in accuracy as context length increases. Additionally, we find that while grouping relevant information generally improves performance, the optimal positioning varies across models. Our findings surface both the advancements and the ongoing challenges in achieving a holistic understanding of long contexts. These can guide future developments in LCLMs and set the stage for creating more robust language models for real-world applications. Code https://github.com/megagonlabs/holobench Benchmark https://hf.co/datasets/megagonlabs/holobench * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Reliable Benchmarking: A Contamination Free, Controllable Evaluation Framework for Multi-step LLM Function CallingSeiji Maekawa, Jackson Hassell, Pouya Pezeshkpour, Tom M. Mitchell 等ICLR 2026 · 被引用 14 次
- Same Content, Different Representations: A Controlled Study for Table QAYue Zhang, Seiji Maekawa, Nikita BhutaniICLR 2026 · 被引用 5 次
- Relational Deep Dive: Error-Aware Queries Over Unstructured DataDaren Chao, Kaiwen Chen, Naiqing Guan, Nick KoudasVLDB 2026 · 被引用 3 次
- DiffBench Meets DiffAgent: End-to-End LLM-Driven Diffusion Acceleration Code GenerationJiajun Jiao, Haowei Zhu, Puyuan Yang, Jianghui Wang 等AAAI 2026 · 被引用 1 次
- Efficient Context Selection for Long-Context QA: No Tuning, No Iteration, Just Adaptive-kChihiro Taguchi, Seiji Maekawa, Nikita BhutaniEMNLP 2025
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationFangyuan Xu, Weijia Shi, Eunsol ChoiICLR 2024 · 被引用 260 次
相关 Paper
- LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context GrowthWeihao Zeng, Yuzhen Huang, Junxian HeICML 2026 · 被引用 11 次
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao 等EMNLP 2024 · 被引用 9 次
- LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs - No Silver Bullet for LC or RAG RoutingKuan Li, Liwen Zhang, Yong Jiang, Pengjun Xie 等ICML 2025
- MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question AnsweringTeng Lin, Yuyu Luo, Honglin Zhang, Jicheng Zhang 等EMNLP 2025 · 被引用 2 次
- RetroLM: Retrieval-Augmented KVs for Long-Context ProcessingKun Luo, Zheng Liu, Shitao Xiao, Jiabei Chen 等AAAI 2026
