Grounding Language in Multi-Perspective Referential Communication
Zineng Tang, Lingjun Mao, Alane Suhr
Abstract
We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspective, which may be different from their own, to both produce and understand references to objects in a scene and the spatial relations between them. We collect a dataset of 2,970 humanwritten referring expressions, each paired with human comprehension judgments, and evaluate the performance of automated models as speakers and listeners paired with human partners, finding that model performance in both reference generation and comprehension lags behind that of pairs of human agents. Finally, we experiment training an open-weight speaker model with evidence of communicative success when paired with a listener, resulting in an improvement from 58.9 to 69.3% in communicative success and even outperforming the strongest proprietary model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c756e96-1839-4b52-9642-8cc8fd6b6b03Cited by top-tier papers4
- LVLMs and Humans Ground Differently in Referential CommunicationPeter Zeng, Weiling Li, Amie J. Paige, Zhengxiang Wang et al.ACL 2026 · 3 citations
- Representational Similarity and Model Behavior in Multi-Agent InteractionYujin Potter, Seun Eisape, Shiyang Lai, Alexander Huth et al.ICML 2026
- Evaluating Model Perception of Color Illusions in Photorealistic ScenesLingjun Mao, Zineng Tang, Alane SuhrCVPR 2025
- LVLMs are Bad at Overhearing Human Referential CommunicationZhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow et al.EMNLP 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du et al.ICLR 2024 · 515 citations
Related papers
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
- PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesMd Mofijul Islam, Alexi Gladstone, Tariq IqbalAAAI 2023 · 10 citations
- Learning Multi-Object Positional Relationships via Emergent CommunicationYicheng Feng, Boshi An, Zongqing LuAAAI 2024 · 4 citations
- Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionZhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong et al.CVPR 2020
- RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4DShuhei Kurita, Naoki Katsura, Eri OnamiICCV 2023 · 26 citations
