Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models
Junjie Wu, Gefei Gu, Yanan Zheng, Dit-Yan Yeung, Arman Cohan
Abstract
Long-context language models (LCLMs) have exhibited impressive capabilities in longcontext understanding tasks. Among these, long-context referencing-a crucial task that requires LCLMs to attribute items of interest to specific parts of long-context data-remains underexplored. To bridge this gap, this paper proposes Referencing Evaluation for Longcontext Language Models (Ref-Long), a novel benchmark designed to assess the long-context referencing capability of LCLMs. Specifically, Ref-Long requires LCLMs to identify the indexes of documents that reference a specific key, emphasizing contextual relationships between the key and the documents over simple retrieval. Based on the task design, we construct three subsets ranging from synthetic to realistic scenarios to form the Ref-Long benchmark. Experimental results of 13 LCLMs reveal significant shortcomings in long-context referencing, even among advanced models like GPT-4o. To further investigate these challenges, we conduct comprehensive analyses, including human evaluations, task format adjustments, fine-tuning experiments, and error analyses, leading to several key insights. Our data and code can be found in https://github. com/wujunjie1998/Ref-Long . * Equal contribution. 1 The term referencing differs from retrieval in that it requires LCLMs to not only retrieve keys from long context, Tell me the indexes of all sections referencing Durant. 2 1 …Anthony and Jeremy Lin work together for New York…
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsMosh Levy, Alon Jacoby, Yoav GoldbergACL 2024 · 77 citations
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 19 citations
- QuerySum: A Multi-Document Query-Focused Summarization Dataset Augmented with Similar Query ClustersYushan Liu, Zili Wang, Ruifeng YuanAAAI 2024 · 14 citations
- Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAMinzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao et al.EMNLP 2024 · 9 citations
Related papers
- L-Eval: Instituting Standardized Evaluation for Long Context Language ModelsChenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao et al.ACL 2024 · 6 citations
- M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun et al.ACL 2024 · 3 citations
- LongCodeU: Benchmarking Long-Context Language Models on Long Code UnderstandingJia Li, Xuyuan Guo, Lei Li, Kechi Zhang et al.ACL 2025
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu et al.ACL 2024
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
