Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding
Biswesh Mohapatra, Manav Nitin Kapadnis, Laurent Romary, Justine Cassell
Abstract
Conversational grounding, vital for building effective dialogue between people and between people and dialogue systems, involves ensuring a mutual understanding of shared information. Despite its importance, there has been limited research on this aspect of conversation in recent years, especially after the advent of Large Language Models (LLMs). Previous studies have highlighted the shortcomings of some pre-trained language models in conversational grounding. However, most testing for conversational grounding capabilities involves human evaluations that are costly and time-consuming. This has led to a lack of testing across multiple models of varying sizes, a critical need given the rapid rate of new model releases. This gap in research becomes more significant considering recent advances in language models, which have led to new emergent capabilities. In this paper, we evaluate the performance of LLMs in various aspects of conversational grounding and analyze why some models perform better than others. We demonstrate a direct correlation between the size of the pre-training dataset, size of the model and conversational grounding abilities, suggesting that they have independently acquired some pragmatic capabilities from larger pre-training datasets. Finally, we propose ways to enhance the capabilities of the models that lag in our tests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df9782f5-449c-4a4f-bcba-2adf1bbf7aa1Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Reference-Centric Models for Grounded Collaborative DialogueDaniel Fried, Justin T. Chiu, Dan KleinEMNLP 2021 · 12 citations
Related papers
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political QuestionsClara Lachenmaier, Judith Sieker, Sina ZarrießACL 2025
- Symbolic Planning and Code Generation for Grounded DialogueJustin T. Chiu, Wenting Zhao, Derek Chen, Saujas Vaduguru et al.EMNLP 2023
- Call for Customized Conversation: Customized Conversation Grounding Persona and KnowledgeYoonna Jang, Jungwoo Lim, Yuna Hur, Dongsuk Oh et al.AAAI 2022 · 47 citations
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
