Refer, Reuse, Reduce: Generating Subsequent References in Visual and Conversational Contexts
Ece Takmaz, Mario Giulianelli, Sandro Pezzelle, Arabella Sinclair, Raquel Fernández
Abstract
Dialogue participants often refer to entities or situations repeatedly within a conversation, which contributes to its cohesiveness. Subsequent references exploit the common ground accumulated by the interlocutors and hence have several interesting properties, namely, they tend to be shorter and reuse expressions that were effective in previous mentions. In this paper, we tackle the generation of first and subsequent references in visually grounded dialogue. We propose a generation model that produces referring utterances grounded in both the visual and the conversational context. To assess the referring effectiveness of its output, we also implement a reference resolution system. Our experiments and analyses show that the model produces better, more effective referring utterances than a model not grounded in the dialogue context, and generates subsequent references that exhibit linguistic patterns akin to humans. Referring utterances extracted from dialogue 1 A: a white fuzzy dog with a wine glass up to his face ; B: I see the wine glass dog ; A: no I don't have the wine glass dog Referring utterances extracted from dialogue 2 C: white dog sitting on something red ; D: yes I have the dog on the red chair ; C: white dog on the red chair
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7642f2f4-1a50-4ea5-8d4c-e688e444cf95Cited by top-tier papers5
- Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and ChallengesBolei Ma, Yuting Li, Wei Zhou, Ziwei Gong et al.ACL 2025 · 28 citations
- Reference-Centric Models for Grounded Collaborative DialogueDaniel Fried, Justin T. Chiu, Dan KleinEMNLP 2021 · 12 citations
- Dealing with Semantic Underspecification in Multimodal NLPSandro PezzelleACL 2023 · 5 citations
- When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue AgentsJiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long et al.ACL 2026 · 5 citations
- Success and Cost Elicit Convention Formation for Efficient CommunicationSaujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel FriedACL 2026 · 3 citations
Builds on1
Related papers
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- Who are you referring to? Coreference resolution in image narrationsArushi Goel, Basura Fernando, Frank Keller, Hakan BilenICCV 2023 · 8 citations
- Learning Reasoning Paths over Semantic Graphs for Video-grounded DialoguesHung Le, Nancy F. Chen, Steven C. H. HoiICLR 2021 · 18 citations
- Whether you can locate or not? Interactive Referring Expression GenerationFulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie WangACM MM 2023 · 6 citations
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang et al.ACM MM 2022 · 2 citations
