Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language Models
Francisco Javier Chiyah Garcia, Alessandro Suglia, Arash Eshghi
摘要
In dialogue, the addressee may initially misunderstand the speaker and respond erroneously, often prompting the speaker to correct the misunderstanding in the next turn with a Third Position Repair (TPR).The ability to process and respond appropriately to such repair sequences is thus crucial in conversational AI systems.In this paper, we first collect, analyse, and publicly release BLOCKWORLD-REPAIRS: a dataset of multi-modal TPR sequences in an instruction-following manipulation task that is, by design, rife with referential ambiguity.We employ this dataset to evaluate several state-ofthe-art Vision and Language Models (VLM) across multiple settings, focusing on their capability to process and accurately respond to TPRs and thus recover from miscommunication.We find that, compared to humans, all models significantly underperform in this task.We then show that VLMs can benefit from specialised losses targeting relevant tokens during fine-tuning, achieving better performance and generalising better to new scenarios.Our results suggest that these models are not yet ready to be deployed in multi-modal collaborative settings where repairs are common, and highlight the need to design training regimes and objectives that facilitate learning from interaction.Our code and data are available at www.github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 被引用 401 次
- TEACh: Task-Driven Embodied Agents That ChatAishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange 等AAAI 2022 · 被引用 251 次
相关 Paper
- TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal ModelsZhenkun Gao, Xuhong Wang, Xin Tan, Yuan XieICLR 2026 · 被引用 1 次
- Understanding is a Two-Way Street: User-Initiated Repair on Agent Responses and Hearing in Conversational InterfacesRobert J. Moore, Sungeun An, Olivia H. MarreseCSCW 2024 · 被引用 3 次
- Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMsClara Lachenmaier, Hannah Bultmann, Sina ZarrießACL 2026
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 被引用 37 次
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier 等EMNLP 2024 · 被引用 5 次
