HELMET: How to Evaluate Long-context Models Effectively and Thoroughly
Howard Yen, Tianyu Gao, Minmin Hou, Ke Ding, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, Danqi Chen
Abstract
There have been many benchmarks for evaluating long-context language models (LCLMs), but developers often rely on synthetic tasks like needlein-a-haystack (NIAH) or arbitrary subsets of tasks. It remains unclear whether they translate to the diverse downstream applications of LCLMs, and the inconsistency further complicates model comparison. We investigate the underlying reasons behind current practices and find that existing benchmarks often provide noisy signals due to low coverage of applications, insufficient lengths, unreliable metrics, and incompatibility with base models. In this work, we present HELMET (How to Evaluate Longcontext Models Effectively and Thoroughly), a comprehensive benchmark encompassing seven diverse, application-centric categories. We also address many issues in previous benchmarks by adding controllable lengths up to 128k tokens, model-based evaluation for reliable metrics, and fewshot prompting for robustly evaluating base models. Consequently, we demonstrate that HELMET offers more reliable and consistent rankings of frontier LCLMs. Through a comprehensive study of 51 LCLMs, we find that (1) synthetic tasks like NIAH are not good predictors of downstream performance; (2) the diverse categories in HELMET exhibit distinct trends and low correlation with each other; and (3) while most LCLMs achieve perfect NIAH scores, open-source models significantly lag behind closed ones when the task requires full-context reasoning or following complex instructions-the gap widens with increased lengths. Finally, we recommend using our RAG tasks for fast model development, as they are easy to run and more predictive of other downstream performance; ultimately, we advocate for a holistic evaluation across diverse tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language ModelsZiyi Yang, Weizhou Shen, Chenliang Li, Ruijun Chen et al.ICLR 2026 · 27 citations
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining DataHaoran Deng, Yingyu Lin, Zhenghao Lin, Xiao Liu et al.ICLR 2026 · 6 citations
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu et al.ICLR 2026 · 5 citations
- HGMem: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational ModelingChulun Zhou, Chunkang Zhang, Guoxin Yu, Fandong Meng et al.ICML 2026 · 4 citations
- AdaSplash-2: Faster Differentiable Sparse AttentionNuno M. T. Gonçalves, Hugo Pitorro, Vlad Niculae, Edoardo Ponti et al.ICML 2026 · 3 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe et al.EMNLP 2022 · 634 citations
Related papers
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMsYuhao Wu, Ming Shan Hee, Zhiqiang Hu, Roy Ka-Wei LeeICLR 2025
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long ContextsYifei Yu, Qian-Wen Zhang, Lingfeng Qiao, Di Yin et al.EMNLP 2025 · 2 citations
- L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao et al.ACL 2024 · 6 citations
- NoLiMa: Long-Context Evaluation Beyond Literal MatchingAli Modarressi, Hanieh Deilamsalehy, Franck Dernoncourt, Trung Bui et al.ICML 2025
- LongSafety: Evaluating Long-Context Safety of Large Language ModelsYida Lu, Jiale Cheng, Zhexin Zhang, Shiyao Cui et al.ACL 2025 · 6 citations
