MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models
Zhongzhan Huang, Guoming Ling, Shanshan Zhong, Hefeng Wu, Liang Lin
Abstract
Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs). However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses. Through extensive experimentation, we discover that existing LCU benchmarks exhibit significant redundancy, which means the inefficiency in evaluation. In this paper, we propose a concise data compression method tailored for longtext data with sparse information characteristics. By pruning the well-known LCU benchmark LongBench, we create MiniLongBench. This benchmark includes only 237 test samples across six major task categories and 21 distinct tasks. Through empirical analysis of over 60 LLMs, MiniLongBench achieves an average evaluation cost reduced to only 4.5% of the original while maintaining an average rank correlation coefficient of 0.97 with Long-Bench results. Therefore, our MiniLongBench, as a low-cost benchmark, holds great potential to substantially drive future research into the LCU capabilities of LLMs. See Github for our code, data and tutorial.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec9d47fd-a2fb-454b-90c8-f868c728b80cCited by top-tier papers3
- AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic WebShanshan Zhong, Kate Shen, Chenyan XiongICML 2026 · 2 citations
- Massive Editing for Large Language Models Based on Dynamic Weight GenerationWentao Wan, Qiqing Lao, Zhiwei Xie, Hefeng Wu et al.ICLR 2026 · 1 citation
- DART: Dual Adaptive Refinement Transfer for Open-Vocabulary Multi-Label RecognitionHaijing Liu, Tao Pu, Hefeng Wu, Keze Wang et al.ACM MM 2025
Builds on29
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando et al.ICML 2023 · 474 citations
Related papers
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
- CMedBench: A Comprehensive Benchmark for Efficient Medical Large Language ModelsShengbo Gao, Jinyang Guo, Lixian Su, Yifu Ding et al.AAAI 2026
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu et al.ACL 2024
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionPeijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li et al.ICML 2025
- LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt CompressionHuiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li et al.ACL 2024 · 59 citations
