Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi
Abstract
In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR).The ability to process and respond appropriately to such repair sequences is thus crucial in conversational AI systems.In this paper, we first collect, analyse, and publicly release BLOCKWORLD-REPAIRS: a dataset of multi-modal TPR sequences in an instruction-following manipulation task that is, by design, rife with referential ambiguity.We employ this dataset to evaluate several state-ofthe-art Vision and Language Models (VLM) across multiple settings, focusing on their capability to process and accurately respond to TPRs and thus recover from miscommunication.We find that, compared to humans, all models significantly underperform in this task.We then show that VLMs can benefit from specialised losses targeting relevant tokens during fine-tuning, achieving better performance and generalising better to new scenarios.Our results suggest that these models are not yet ready to be deployed in multi-modal collaborative settings where repairs are common, and highlight the need to design training regimes and objectives that facilitate learning from interaction.Our code and data are available at www.github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a68edc9-3349-4fbf-9488-252650f3c994Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange et al.AAAI 2022 · 251 citations
Related papers
- TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal ModelsZhenkun Gao, Xuhong Wang, Xin Tan, Yuan XieICLR 2026 · 1 citation
- Understanding is a Two-Way Street: User-Initiated Repair on Agent Responses and Hearing in Conversational InterfacesRobert J. Moore, Sungeun An, Olivia H. MarreseCSCW 2024 · 3 citations
- Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMsClara Lachenmaier, Hannah Bultmann, Sina ZarrießACL 2026
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 37 citations
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
