Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation Tasks
Amit Parekh, Nikolas Vitsakis, Alessandro Suglia, Ioannis Konstas
Abstract
Evaluating the generalisation capabilities of multimodal models based solely on their performance on out-of-distribution data fails to capture their true robustness. This work introduces a comprehensive evaluation framework that systematically examines the role of instructions and inputs in the generalisation abilities of such models, considering architectural design, input perturbations across language and vision modalities, and increased task complexity. The proposed framework uncovers the resilience of multimodal models to extreme instruction perturbations and their vulnerability to observational changes, raising concerns about overfitting to spurious correlations. By employing this evaluation framework on current Transformerbased multimodal models for robotic manipulation tasks, we uncover limitations and suggest future advancements should focus on architectural and training innovations that better integrate multimodal inputs, enhancing a model's generalisation prowess by prioritising sensitivity to input content over incidental correlations. 1 L1 L2 L3 L4 (a) Trained on Original; Evaluated on Original
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual LearningZiqi Jia, Anmin Wang, Xiaoyang Qu, Xiaowen Yang et al.ACL 2025
- Repairs in a Block World: A New Benchmark for Handling User Corrections with Multi-Modal Language ModelsFrancisco Javier Chiyah Garcia, Alessandro Suglia, Arash EshghiEMNLP 2024
Builds on14
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action ModelsSenyu Fei, Siyin Wang, Junhao Shi, Zihao Dai et al.CVPR 2026
- On the generalization capacity of neural networks during generic multimodal reasoningTakuya Ito, Soham Dan, Mattia Rigotti, James R. Kozloski et al.ICLR 2024 · 4 citations
- VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic ManipulationZhijie Wang, Zhehua Zhou, Jiayang Song, Yuheng Huang et al.FSE 2025 · 6 citations
- MMIFEvol: Towards Evolutionary Multimodal Instruction FollowingHaoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen et al.AAAI 2026 · 1 citation
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsBorong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang et al.ICML 2026 · 25 citations
