RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation
Sanghwan Jang, Minjin Jeon, Minsoo Kim, Seong Jin Choi, Dongha Kim, Hwanjo Yu
Abstract
Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SAFE: Multitask Failure Detection for Vision-Language-Action ModelsQiao Gu, Yuanliang Ju, Shengxiang Sun, Igor Gilitschenski et al.NeurIPS 2025 · 103 citations
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language ModelsGuo Chen, Zhiqi Li, Shihao Wang, Jindong Jiang et al.NeurIPS 2025 · 69 citations
- Retrieval-Augmented Embodied AgentsYichen Zhu, Zhicai Ou, Xiaofeng Mou, Jian TangCVPR 2024 · 9 citations
Related papers
- FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action AdaptationDuc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho et al.ICML 2026 · 3 citations
- SimpleVLA-RL: Scaling VLA Training via Reinforcement LearningHaozhan Li, Yuxin Zuo, Jiale Yu, Yuhao Zhang et al.ICLR 2026 · 170 citations
- MetaVLA: Unified Meta Co-Training for Efficient Embodied AdaptationChen Li, Zhantao Yang, Han Zhang, Fangyi Chen et al.ICLR 2026 · 2 citations
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action ModelsXiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang et al.ICLR 2026 · 59 citations
- VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token CachingSiyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu et al.NeurIPS 2025 · 95 citations
