DRInQ: Evaluating Conversational Implicature with Controlled Context Variation
Hirona Jacqueline Arai, Xiang Ren
摘要
Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Although recent large language models (LLMs) exhibit strong conversational fluency, they remain unreliable when interpretation depends on reasoning that integrates social and contextual cues, a process rarely articulated in text. We introduce DRinQ, a benchmark for evaluating pragmatic reasoning about conversational implicature in question utterances, designed to isolate pragmatic variation while holding each question's surface form fixed. To support scalable evaluation, we propose a semi-automated pipeline that produces question-context-interpretation instances with systematic variation. Across evaluations, we find a consistent generation-inference asymmetry: while state-of-the-art models can generate plausible pragmatic scenarios when guided, they often fail to recover the intended implication at inference time. For smaller models, structured prompting improves alignment with human judgments. A comparative writing study further reveals complementary strengths: human authors tend to produce safer, predictable contexts, whereas models generate varied scenarios with interpretations that sometimes exceed contextual support. These findings highlight persistent challenges in modeling conversational implicature and motivate more contextsensitive evaluation frameworks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMsLaura Ruis, Akbir Khan, Stella Biderman, Sara Hooker 等NeurIPS 2023 · 被引用 87 次
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 被引用 52 次
- How do you Converse with an Analytical Chatbot? Revisiting Gricean Maxims for Designing Analytical Conversational BehaviorVidya Setlur, Melanie ToryCHI 2022 · 被引用 46 次
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko 等ACL 2023 · 被引用 45 次
相关 Paper
- FLUID QA: A Multilingual Benchmark for Figurative Language Usage in Dialogue across English, Chinese, and KoreanSeoyoon Park, Hyeji Choi, Minseon Kim, Subin An 等EMNLP 2025 · 被引用 1 次
- DyVal: Dynamic Evaluation of Large Language Models for Reasoning TasksKaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong 等ICLR 2024 · 被引用 92 次
- PRISM: Probing Reasoning, Instruction, and Source Memory in LLM HallucinationsYuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang 等ACL 2026
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen 等NeurIPS 2024 · 被引用 215 次
- Beyond Facts: Evaluating Intent Hallucination in Large Language ModelsYijie Hao, Haofei Yu, Jiaxuan YouACL 2025
