On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts
Linlu Qiu, Cedegao E. Zhang, Joshua B. Tenenbaum, Yoon Kim, Roger P. Levy
Abstract
Language use is shaped by pragmatics-i.e., reasoning about communicative goals and norms in context. As language models (LMs) are increasingly used as conversational agents, it becomes ever more important to understand their pragmatic reasoning abilities. We propose an evaluation framework derived from Wavelength, a popular communication game where a speaker and a listener communicate about a broad range of concepts in a granular manner. We study a range of LMs on both language comprehension and language production using direct and Chain-of-Thought (CoT) prompting, and further explore a Rational Speech Act (RSA) approach to incorporating Bayesian pragmatic reasoning into LM inference. We find that state-of-the-art LMs, but not smaller ones, achieve strong performance on language comprehension, obtaining similar-to-human accuracy and exhibiting high correlations with human judgments even without CoT prompting or RSA. On language production, CoT can outperform direct prompting, and using RSA provides significant improvements over both approaches. Our study helps identify the strengths and limitations in LMs' pragmatic reasoning abilities and demonstrates the potential for improving them with RSA, opening up future avenues for understanding conceptual representation, language understanding, and social reasoning in LMs and humans. 1 Left Concept (0) Target Value Right Concept (100) Human-written Clues Chosen Clue Human Mean Deep thought 10 Shallow thought Evolution, Solving complex problems, Chess, Einstein, Meditation, Quantum mechanics Solv. complex prob.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Evaluating Language Models' Evaluations of GamesKatherine M. Collins, Cedegao E. Zhang, Graham Todd, Lance Ying et al.ICLR 2026 · 5 citations
- Cognitive models can reveal interpretable value trade-offs in language modelsSonia Krishna Murthy, Rosie Zhao, Jennifer Hu, Sham M. Kakade et al.ICLR 2026 · 2 citations
- DRInQ: Evaluating Conversational Implicature with Controlled Context VariationHirona Jacqueline Arai, Xiang RenACL 2026
Builds on8
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMsLaura Ruis, Akbir Khan, Stella Biderman, Sara Hooker et al.NeurIPS 2023 · 87 citations
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko et al.ACL 2023 · 45 citations
- How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?Ryan Liu, Theodore R. Sumers, Ishita Dasgupta, Thomas L. GriffithsICML 2024 · 33 citations
Related papers
- Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than HelpsHaibo Jin, Peiyan Zhang, Man Luo, Haohan WangNeurIPS 2025 · 1 citation
- Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn DialogLautaro Estienne, Gabriel Ben Zenou, Nona Naderi, Jackie CK Cheung et al.EMNLP 2025
- ReaGEN: Adaptive Generation of Structured Chains-of-Thought for Efficient Multimodal ReasoningRuiqing Tian, Mohan Sai Singamsetti, Di Niu, Bahador RashidiCVPR 2026
- Rethinking Pragmatics in Large Language Models: Towards Open-Ended Evaluation and Preference TuningShengguang Wu, Shusheng Yang, Zhenglun Chen, Qi SuEMNLP 2024 · 2 citations
- Understanding Chain-of-Thought in LLMs through Information TheoryJean-Francois Ton, Muhammad Faaiz Taufiq, Yang LiuICML 2025
