LVLMs and Humans Ground Differently in Referential Communication
Peter Zeng, Weiling Li, Amie J. Paige, Zhengxiang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory J. Zelinsky, Susan Brennan, Owen Rambow
Abstract
For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a critical deficit: an inability to model common ground. We present a referential communication experiment with a factorial design involving director-matcher pairs (human-human, human-AI, AI-human, and AI-AI) that interact with multiple turns in repeated rounds to match pictures of objects not associated with any obvious lexicalized labels. We show that LVLMs cannot interactively generate and resolve referring expressions in a way that enables smooth communication, a crucial skill that underlies human language use. We release our corpus of 356 dialogues (89 pairs over 4 rounds each) along with the online pipeline for data collection and the tools for analyzing accuracy, efficiency, and lexical overlap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46a905dc-37bf-4d9f-9f2d-46b84ebea852Builds on5
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Navigating Rifts in Human-LLM Grounding: Study and BenchmarkOmar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney et al.ACL 2025 · 21 citations
- Understanding Common Ground Misalignment in Goal-Oriented Dialog: A Case-Study with Ubuntu Chat LogsRupak Sarkar, Neha Srikanth, Taylor Pellegrin, Rachel Rudinger et al.ACL 2025 · 3 citations
- Grounding Language in Multi-Perspective Referential CommunicationZineng Tang, Lingjun Mao, Alane SuhrEMNLP 2024 · 1 citation
- LVLMs are Bad at Overhearing Human Referential CommunicationZhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow et al.EMNLP 2025
Related papers
- Learning Multi-Object Positional Relationships via Emergent CommunicationYicheng Feng, Boshi An, Zongqing LuAAAI 2024 · 4 citations
- Emergent Communication of GeneralizationsJesse Mu, Noah D. GoodmanNeurIPS 2021 · 60 citations
- Reference-Centric Models for Grounded Collaborative DialogueDaniel Fried, Justin T. Chiu, Dan KleinEMNLP 2021 · 12 citations
- Learning "Partner-Aware" Collaborators in Multi-Party CollaborationAbhijnan Nath, Nikhil KrishnaswamyNeurIPS 2025 · 2 citations
- An Annotated Corpus of Reference Resolution for Interpreting Common GroundingTakuma Udagawa, Akiko AizawaAAAI 2020 · 10 citations
