Lune

ACL2025顶会

Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models

Junjie Wu, Gefei Gu, Yanan Zheng, Dit-Yan Yeung, Arman Cohan

2025年份
4被引次数
1顶会引用

摘要

Long-context language models (LCLMs) have exhibited impressive capabilities in longcontext understanding tasks. Among these, long-context referencing-a crucial task that requires LCLMs to attribute items of interest to specific parts of long-context data-remains underexplored. To bridge this gap, this paper proposes Referencing Evaluation for Longcontext Language Models (Ref-Long), a novel benchmark designed to assess the long-context referencing capability of LCLMs. Specifically, Ref-Long requires LCLMs to identify the indexes of documents that reference a specific key, emphasizing contextual relationships between the key and the documents over simple retrieval. Based on the task design, we construct three subsets ranging from synthetic to realistic scenarios to form the Ref-Long benchmark. Experimental results of 13 LCLMs reveal significant shortcomings in long-context referencing, even among advanced models like GPT-4o. To further investigate these challenges, we conduct comprehensive analyses, including human evaluations, task format adjustments, fine-tuning experiments, and error analyses, leading to several key insights. Our data and code can be found in https://github. com/wujunjie1998/Ref-Long . * Equal contribution. 1 The term referencing differs from retrieval in that it requires LCLMs to not only retrieve keys from long context, Tell me the indexes of all sections referencing Durant. 2 1 …Anthony and Jeremy Lin work together for New York…

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖