SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Haozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang, Yang Zhaohui, Kaiyan Zhang, Xuekai Zhu, Yuchen Zhang, Tianxing Chen, Ganqu Cui, Dehui Wang, Dingxiang Luo
Abstract
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for robotic manipulation. Despite substantial progress enabled by large-scale pretraining and supervised fine-tuning (SFT), these models face two fundamental challenges: (i) the scarcity and high cost of large-scale robotic trajectories required for SFT scaling, and (ii) limited generalization to tasks under distribution shift. To overcome these limitations, we explore reinforcement learning (RL) as a pathway to scaling VLA training beyond limited datasets. Inspired by LLM breakthroughs where RL with outcome rewards enhances step-by-step reasoning, we ask: Can outcome-driven RL improve long-horizon step-by-step action planning of VLA? In this work, we introduce SimpleVLA-RL, an efficient RL framework tailored for VLA models. Building upon veRL, we introduce VLA-specific trajectory sampling, scalable parallelization, multi-environment rendering, and optimized loss computation. Applied to OpenVLA-OFT, SimpleVLA-RL achieves 99% of SoTA performance on LIBERO and 80% relative improvement on RoboTwin 1.0&2.0, outperforming with our proposed exploration-enhancing strategies. SimpleVLA-RL reduces dependence on large-scale data, enables robust generalization, and remarkably surpasses SFT in real-world tasks. Moreover, we identify a novel phenomenon "pushcut'' during RL training, wherein the policy discovers unseen patterns beyond those seen in previous training process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c954ecb-60a2-4052-8a4e-c2b8606b552aCited by top-tier papers23
- WMPO: World Model-based Policy Optimization for Vision-Language-Action ModelsFangqi Zhu, Zhengyang Yan, Zicong Hong, Quanxin Shou et al.ICLR 2026 · 64 citations
- RLinf: Flexible and Efficient Large-Scale Reinforcement Learning via Macro-to-Micro Flow TransformationChao Yu, Yuanqing Wang, Zhen Guo, Hao Lin et al.OSDI 2026 · 63 citations
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelYanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen et al.ICML 2026 · 29 citations
- SRPO: Self-Referential Policy Optimization for Vision-Language-Action ModelsSenyu Fei, Siyin Wang, Li Ji, Ao Li et al.CVPR 2026 · 28 citations
- Mixture of Horizons in Action ChunkingDong Jing, Gang Wang, Jiaqi Liu, Weiliang Tang et al.ICML 2026 · 24 citations
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu et al.NeurIPS 2025 · 181 citations
Related papers
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language ModelsZirui Song, Guangxian Ouyang, Mingzhe Li, Yuheng Ji et al.AAAI 2026 · 21 citations
- On-the-Fly VLA Adaptation via Test-Time Reinforcement LearningChangyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang et al.ACL 2026 · 7 citations
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RLWenli Xiao, Haotian Lin, Andy Peng, Haoru Xue et al.ICLR 2026 · 84 citations
- What Can RL Bring to VLA Generalization? An Empirical StudyJijia Liu, Feng Gao, Bingwen Wei, Xinlei Chen et al.NeurIPS 2025 · 120 citations
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action ModelBing Hu, Zaijing Li, Rui Shao, Junda Chen et al.ICML 2026 · 4 citations
