LooGLE: Can Long-Context Language Models Understand Long Contexts?
Jiaqi Li, Mengmeng Wang, Zilong Zheng, Muhan Zhang
摘要
Large language models (LLMs) are typically limited to processing texts within contextwindow size, which has spurred significant research efforts into enhancing LLMs' longcontext understanding as well as developing high-quality benchmarks to evaluate the ability. However, prior datasets suffer from shortcomings like short length compared to the context window of modern LLMs; outdated documents that might have data leakage problems; and an emphasis on short dependency tasks only. In this paper, we present ooGLE , a Long Context Generic Language Evaluation benchmark. It features documents post-2022, with over 24,000 tokens per document and 6,000 newly generated questions spanning varying dependency ranges in diverse domains. Human annotators meticulously crafted over 1,100 high-quality question-answer (QA) pairs with thorough cross-validation for a most precise assessment of LLMs' long dependency capabilities. We conduct a comprehensive evaluation of representative LLMs on ooGLE . The results indicate that most LLMs have shockingly bad long context ability and fail to capture long dependencies in the context, even when their context window size is enough to fit the entire document. Our results shed light on enhancing the "true long-context understanding" ability of LLMs instead of merely enlarging their context window.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang 等NeurIPS 2024 · 被引用 151 次
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 被引用 77 次
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu 等ASPLOS 2026 · 被引用 18 次
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic AnchorsXin Liu, Runsong Zhao, Pengcheng Huang, Xinyu Liu 等ICLR 2026 · 被引用 16 次
- Statically Contextualizing Large Language Models with Typed HolesAndrew Blinn, Xiang Li, June Hyung Kim, Cyrus OmarOOPSLA 2024 · 被引用 10 次
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
相关 Paper
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu 等ACL 2024
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao 等EMNLP 2024 · 被引用 9 次
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensCunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu 等ICLR 2025
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context MultitasksYushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng 等ACL 2025
