Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMs
Amber Shore, Russell Scheinberg, Ameeta Agrawal, So Young Lee
Abstract
Large Language Models (LLMs) are intended to reflect human linguistic competencies. But humans have access to a broad and embodied context, which is key in detecting and resolving linguistic ambiguities, even in isolated text spans. A foundational case of semantic ambiguity is found in the task of coreference resolution: how is a pronoun related to an earlier person mention? This capability is implicit in nearly every downstream task, and the presence of ambiguity at this level can alter performance significantly. We show that LLMs can achieve good performance with minimal prompting in both coreference disambiguation and the detection of ambiguity in coreference, however, they cannot do both at the same time. We present the CORRECT-DETECT trade-off: though models have both capabilities and deploy them implicitly, successful performance balancing these two abilities remains elusive. Unambiguous Kimberly told the aunt that she contacted the granddaughter. Who contacted the granddaughter? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Kimberly told the aunt that the granddaughter trusted her. Who did the granddaughter trust? Ambiguous Mostly ambiguous Overdetects ambiguity Detect Ambig
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Is There a Trade-Off Between Fairness and Accuracy? A Perspective Using Mismatched Hypothesis TestingSanghamitra Dutta, Dennis Wei, Hazar Yueksel, Pin-Yu Chen et al.ICML 2020 · 171 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
- A Survey on Asking Clarification Questions Datasets in Conversational SystemsHossein A. Rahmani, Xi Wang, Yue Feng, Qiang Zhang et al.ACL 2023 · 7 citations
- Aligning Language Models to Explicitly Handle AmbiguityHyuhng Joon Kim, Youna Kim, Cheonbok Park, Junyeob Kim et al.EMNLP 2024 · 3 citations
Related papers
- Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMsZhuoxuan Zhang, Jinhao Duan, Edward Kim, Kaidi XuEMNLP 2025 · 1 citation
- EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language ModelsJiacheng Huang, Ning Yu, Xiaoyin YiAAAI 2026
- Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesYu Wang, Emmanuele Chersoni, Chu-Ren HuangACL 2026
- AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentAnastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev et al.ACL 2025
- Probing for Referential Information in Language ModelsIonut-Teodor Sorodoc, Kristina Gulordava, Gemma BoledaACL 2020 · 31 citations
