OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
Fanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You, Junming Zhao, Yang Gao
Abstract
General-purpose robots capable of performing diverse tasks require synergistic reasoning and acting capabilities. However, recent dual-system approaches, which separate high-level reasoning from low-level acting, often suffer from challenges such as limited mutual understanding of capabilities between systems and latency issues. This paper introduces OneTwoVLA, a single unified vision-languageaction model that can perform both acting (System One) and reasoning (System Two). Crucially, OneTwoVLA adaptively switches between two modes: explicitly reasoning at critical moments during task execution, and generating actions based on the most recent reasoning at other times. To further unlock OneTwoVLA's reasoning and generalization capabilities, we design a scalable pipeline for synthesizing embodied reasoning-centric vision-language data, used for co-training with robot data. We validate OneTwoVLA's effectiveness through extensive experiments, highlighting its superior performance across four key capabilities: longhorizon task planning, error detection and recovery, natural human-robot interaction, and generalizable visual grounding, enabling the model to perform longhorizon, highly dexterous manipulation tasks such as making hotpot or mixing cocktails. Project page: https://one-two-vla.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1abd7df-fbd8-4661-bc1b-0d3f519fbf5cCited by top-tier papers20
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World KnowledgeWenyao Zhang, Hongsi Liu, Zekun Qi, Yunnan Wang et al.NeurIPS 2025 · 244 citations
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action ModelsBorong Zhang, Jiahao Li, Jiachen Shen, Yuhao Zhang et al.ICML 2026 · 25 citations
- Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive ReasoningZhenghao Peng, Wenhao Ding, Yurong You, Yuxiao Chen et al.CVPR 2026 · 25 citations
- AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action ModelsXiaoqi Li, Muhe Cai, Jiadong Xu, Juan Zhu et al.CVPR 2026 · 18 citations
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of ThoughtYao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang et al.NeurIPS 2023 · 453 citations
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu et al.ICLR 2024 · 375 citations
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language ModelsSiddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang et al.ICML 2024 · 306 citations
Related papers
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent PlanningChi-Pin Huang, Yueh-Hua Wu, Min-Hung Chen, Yu-Chiang Frank Wang et al.NeurIPS 2025 · 179 citations
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language NavigationMeng Wei, Chenyang Wan, Jiaqi Peng, Xiqian Yu et al.ICLR 2026 · 77 citations
- SemanticVLA: Towards Semantic Reasoning over Action Memorization via Synergistic Explicit Trace and Latent Action PlanningFei Ni, Zhuo Chen, Yifu Yuan, Zibin Dong et al.CVPR 2026
- Fast-in-Slow: A Dual-System VLA Model Unifying Fast Manipulation within Slow ReasoningHao Chen, Jiaming Liu, Chenyang Gu, Zhuoyang Liu et al.NeurIPS 2025 · 74 citations
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought ReasoningQuanxin Shou, Fangqi Zhu, Shuang Chen, Puxin Yan et al.ICML 2026 · 7 citations
