NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng, Guangsheng Bao, Xiangkun Hu, Zheng Zhang, Qian Wang, Yue Zhang
摘要
Recent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this gap, we introduce NovelQA, a benchmark tailored for evaluating LLMs with complex, extended narratives. Constructed from English novels, NovelQA offers a unique blend of complexity, length, and narrative coherence, making it an ideal tool for assessing deep textual understanding in LLMs. This paper details the design and construction of NovelQA, focusing on its comprehensive manual annotation process and the variety of question types aimed at evaluating nuanced comprehension. Our evaluation of long-context LLMs on NovelQA reveals significant insights into their strengths and weaknesses. Notably, the models struggle with multi-hop reasoning, detail-oriented questions, and handling extremely long inputs, with average lengths exceeding 200,000 tokens. Results highlight the need for substantial advancements in LLMs to enhance their long-context comprehension and contribute effectively to computational literary analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading ComprehensionRui Li, Zeyu Zhang, Xiaohe Bo, Zihang Tian 等NeurIPS 2025 · 被引用 29 次
- The Pensieve Paradigm: Stateful Language Models Mastering Their Own ContextXiaoyuan Liu, Tian Liang, Dongyang Ma, Deyu Zhou 等ICLR 2026 · 被引用 8 次
- LiteraryQA: Towards Effective Evaluation of Long-document Narrative QATommaso Bonomo, Luca Gioffré, Roberto NavigliEMNLP 2025 · 被引用 1 次
- RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement LearningKun Li, Yunxiang Li, Tianhua Zhang, Hongyin Luo 等EMNLP 2025 · 被引用 1 次
- Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang 等ACL 2026
它引用的顶会 Paper13
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye 等NeurIPS 2023 · 被引用 470 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
相关 Paper
- Marathon: A Race Through the Realm of Long Context with Large Language ModelsLei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang 等ACL 2024
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu 等ACL 2024
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao 等EMNLP 2024 · 被引用 9 次
- MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail KnowledgeJie He, Nan Hu, Wanqiu Long, Jiaoyan Chen 等ACL 2026 · 被引用 1 次
- LooGLE: Can Long-Context Language Models Understand Long Contexts?Jiaqi Li, Mengmeng Wang, Zilong Zheng, Muhan ZhangACL 2024 · 被引用 32 次
