M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun, Liangyou Li, Yuxin Jiang, Lifeng Shang, Qun Liu, Kam-Fai Wong
摘要
Managing long sequences has become an important and necessary feature for large language models (LLMs). However, assessing their ability to handle long contexts remains a challenge. This paper introduces M 4 LE, a Multi-ability, Multi-range, Multi-task, Multi-domain benchmark for Long-context Evaluation. It encompasses 36 NLP datasets, covering 11 types of tasks and 12 domains, providing a comprehensive test bed. To address the lack of tasks featuring naturally long sequences, we propose an automatic approach to convert short-sequence tasks into long-sequence scenarios. These scenarios evaluate LLMs' long-context understanding across five key abilities: understanding of single or multiple relevant spans in long contexts based on explicit or semantic hints, and global context understanding. This automatic approach allows us to create instances evenly distributed from 1k to 8k input length. 1 Our evaluation of 11 prominent LLMs reveals that 1) Current LLMs struggle to understand long context, particularly when tasks require multiple-span attention. 2) Semantic retrieval is more difficult for competent LLMs. 3) Models fine-tuned on longer text with position interpolation have comparable performance to those using Neural Tangent Kernel (NTK) aware scaling methods without fine-tuning. We make our benchmark publicly available to encourage future research in this challenging area 2 . CHAPTER I On a hill by the Mississippi … For months no male emerged from the mass. … CHAPTER II It was a frail and blue and lonely Carol who…. … CHAPTER X …said Kennicott, as he unpacked his suit-case. … Explicit Single-Span Semantic Single-Span Explicit Multi-Span Semantic Multi-Span Global Span CHAPTER I On a hill by the Mississippi … For months no male emerged from the mass. … CHAPTER II It was a frail and blue and lonely Carol who…. … CHAPTER X …said Kennicott, as he unpacked his suit-case. … Summarize CHAPTER I. CHAPTER I On a hill by the Mississippi … For months no male emerged from the mass. … CHAPTER II It was a frail and blue and lonely Carol who…. … CHAPTER X …said Kennicott, as he unpacked his suit-case. … Summarize the chapter about Carol. Summarize the first and the last chapters. Summarize the chapters about Carol as well as Kennicott. CHAPTER I On a hill by the Mississippi … For months no male emerged from the mass. … CHAPTER II It was a frail and blue and lonely Carol who…. … CHAPTER X …said Kennicott, as he unpacked his suit-case. … Summarize the whole article. CHAPTER I On a hill by the Mississippi … For months no male emerged from the mass. … CHAPTER II It was a frail and blue and lonely Carol who…. … CHAPTER X …said Kennicott, as he unpacked his suit-case. …
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context ScenariosXiaodong Wu, Minhao Wang, Yichen Liu, Xiaoming Shi 等ACL 2025 · 被引用 22 次
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 被引用 19 次
- MDCure: A Scalable Pipeline for Multi-Document Instruction-FollowingGabrielle Kaili-May Liu, Bowen Shi, Avi Caciularu, Idan Szpektor 等ACL 2025 · 被引用 13 次
- L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao 等ACL 2024 · 被引用 6 次
- KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM EvaluationNikita Tatarinov, Vidhyakshaya Kannan, Haricharana Srinivasa, Arnav Raj 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu 等ACL 2024
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu 等ACL 2024 · 被引用 94 次
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao 等EMNLP 2024 · 被引用 9 次
- MMReD: a Cross-Modal Benchmark for Dense Context ReasoningMaxim Kurkin, Boris Shirokikh, Irina Abdullaeva, Viktoriia Chekalina 等ICLR 2026
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context MultitasksYushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng 等ACL 2025
