Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization
Jiaming Zhou, Ke Ye, Jiayi Liu, Teli Ma, Zifan Wang, Ronghe Qiu, Kun-Yu Lin, Zhilin Zhao, Junwei Liang
Abstract
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this gap, we introduce AGNOSTOS, a novel simulation benchmark designed to rigorously evaluate zeroshot cross-task generalization in manipulation. AGNOSTOS comprises 23 unseen manipulation tasks for test, which are distinct from common training task distributions, and incorporates two levels of generalization difficulty to assess robustness. Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks. To overcome this limitation, we propose Cross-Task In-Context Manipulation (X-ICM), a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks. Additionally, we introduce a dynamics-guided sample selection strategy that identifies relevant demonstrations by capturing cross-task dynamics. On AGNOSTOS, X-ICM significantly improves zero-shot cross-task generalization performance over leading VLA models, achieving improvements of 6.0% over π 0 [1] and 7.9% over VoxPoser [2]. We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation. Project page: https://jiaming- zhou.github.io/AGNOSTOS/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82c6ddaa-f08a-4944-9ca9-6a2cc7a47a2dCited by top-tier papers6
- VLANeXt: Recipes for Building Strong VLA ModelsXiao-Ming Wu, Bin Fan, Kang Liao, Jian-Jian Jiang et al.ICML 2026 · 10 citations
- Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action ModelsHanxin Zhang, Mingshuo Xu, Abdulqader Dhafer, Shigang Yue et al.ICML 2026 · 2 citations
- Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human DemonstrationsWei-Jin Huang, Yueyi Zhang, Yi-Lin Wei, Zhi-Wei Xia et al.CVPR 2026
- Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional ShiftsXueyang Zhou, Yangming Xu, Guiyao Tie, Chaoran Hu et al.ICML 2026
- Motion Dynamics Learning for Few-Shot Embodied AdaptationSibo He, Weiying Xie, Daixun Li, Junhao Zhong et al.ICML 2026
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic ManipulationXitie Zhang, Aming WU, Yahong HanICML 2026
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsXiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang et al.ICLR 2026 · 59 citations
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning TasksShiduo Zhang, Zhe Xu, Peiju Liu, Xiaopeng Yu et al.ICCV 2025 · 12 citations
- VIMA: Robot Manipulation with Multimodal PromptsYunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang et al.ICML 2023 · 80 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
