On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning
Changyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang, James Chenhao Liang, Wenhao Yang, Renjing Xu, Qifan Wang, Dongfang Liu, Cheng Han
Abstract
Vision-Language-Action (VLA) models have recently emerged as a powerful paradigm for general-purpose robot learning, enabling agents to map visual observations and naturallanguage instructions into executable robotic actions. Though popular, they are primarily trained via supervised fine-tuning or trainingtime reinforcement learning, requiring explicit fine-tuning phases, human interventions, or controlled data collection. Consequently, existing methods remain unsuitable for challenging simulated-or physical-world deployments, where robots must respond autonomously and flexibly to evolving environments. To address this limitation, we introduce a Test-Time Reinforcement Learning for VLAs (TT-VLA), a framework that enables on-the-fly policy adaptation during inference. TT-VLA formulates a dense reward mechanism that leverages stepby-step task-progress signals to refine action policies during test time while preserving the SFT/RL-trained priors, making it an effective supplement to current VLA models. Empirical results show that our approach enhances overall adaptability, stability, and task success in dynamic, previously unseen scenarios under simulated and real-world settings. We believe TT-VLA offers a principled step toward selfimproving, deployment-ready VLAs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a80c9cd-bed3-41ea-9a14-1b9c672d782fBuilds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- TTT++: When Does Self-Supervised Test-Time Training Fail or Thrive?Yuejiang Liu, Parth Kothari, Bastien van Delft, Baptiste Bellot-Gurlet et al.NeurIPS 2021 · 469 citations
Related papers
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action ModelsHyeonbeom Choi, Daechul Ahn, Youhan Lee, Taewook Kang et al.ICML 2026 · 2 citations
- SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningHaozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang et al.ICLR 2026 · 170 citations
- AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action ModelsXiaoqi Li, Muhe Cai, Jiadong Xu, Juan Zhu et al.CVPR 2026 · 18 citations
- TTRV: Test-Time Reinforcement Learning for Vision Language ModelsAkshit Singh, Shyam Marjit, Wei Lin, Paul Gavrikov et al.CVPR 2026 · 8 citations
- A Generalist Pair-wise Progress Critic Model for Vision-Language-Action RobotsQi Zhang, shaopeng zhai, Shengzhe Zhang, Litao Liu et al.ICML 2026
