Lune

EMNLP2025顶会

Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs

Amber Shore, Russell Scheinberg, Ameeta Agrawal, So Young Lee

2025年份

摘要

Large Language Models (LLMs) are intended to reflect human linguistic competencies. But humans have access to a broad and embodied context, which is key in detecting and resolving linguistic ambiguities, even in isolated text spans. A foundational case of semantic ambiguity is found in the task of coreference resolution: how is a pronoun related to an earlier person mention? This capability is implicit in nearly every downstream task, and the presence of ambiguity at this level can alter performance significantly. We show that LLMs can achieve good performance with minimal prompting in both coreference disambiguation and the detection of ambiguity in coreference, however, they cannot do both at the same time. We present the CORRECT-DETECT trade-off: though models have both capabilities and deploy them implicitly, successful performance balancing these two abilities remains elusive. Unambiguous Kimberly told the aunt that she contacted the granddaughter. Who contacted the granddaughter? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Mostly ambiguous Overdetects ambiguity Detect Ambig

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 783620b0-0ff9-4e02-9465-cf757d0265cc

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖