Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
Amber Shore, Russell Scheinberg, Ameeta Agrawal, So Young Lee
摘要
Large Language Models (LLMs) are intended to reflect human linguistic competencies. But humans have access to a broad and embodied context, which is key in detecting and resolving linguistic ambiguities, even in isolated text spans. A foundational case of semantic ambiguity is found in the task of coreference resolution: how is a pronoun related to an earlier person mention? This capability is implicit in nearly every downstream task, and the presence of ambiguity at this level can alter performance significantly. We show that LLMs can achieve good performance with minimal prompting in both coreference disambiguation and the detection of ambiguity in coreference, however, they cannot do both at the same time. We present the CORRECT-DETECT trade-off: though models have both capabilities and deploy them implicitly, successful performance balancing these two abilities remains elusive. Unambiguous Kimberly told the aunt that she contacted the granddaughter. Who contacted the granddaughter? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Mostly ambiguous Overdetects ambiguity Detect Ambig
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Is There a Trade-Off Between Fairness and Accuracy? A Perspective Using Mismatched Hypothesis TestingSanghamitra Dutta, Dennis Wei, Hazar Yueksel, Pin-Yu Chen 等ICML 2020 · 被引用 171 次
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr 等EMNLP 2023 · 被引用 35 次
- A Survey on Asking Clarification Questions Datasets in Conversational SystemsHossein A. Rahmani, Xi Wang, Yue Feng, Qiang Zhang 等ACL 2023 · 被引用 7 次
- Aligning Language Models to Explicitly Handle AmbiguityHyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim 等EMNLP 2024 · 被引用 3 次
相关 Paper
- Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMsZhuoxuan Zhang, Jinhao Duan, Edward Kim, Kaidi XuEMNLP 2025 · 被引用 1 次
- EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language ModelsJiacheng Huang, Ning Yu, Xiaoyin YiAAAI 2026
- Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesYu Wang, Emmanuele Chersoni, Chu-Ren HuangACL 2026
- AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentAnastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev 等ACL 2025
- Probing for Referential Information in Language ModelsIonut-Teodor Sorodoc, Kristina Gulordava, Gemma BoledaACL 2020 · 被引用 31 次
