VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
Christos Ziakas, Alessandra Russo
Abstract
Vision–Language Models (VLMs) show promise as zero-shot goal-conditioned value functions, but their frozen pre-trained representations limit generalization and temporal reasoning. We introduce VITA, a zero-shot value function learning method that enhances both capabilities via test-time adaptation. At inference, a lightweight adaptation module is updated via a gradient step on a meta-learned self-supervised loss, such that each test-time update improves value estimation. By updating sequentially over a trajectory, VITA encodes history into its parameters, addressing the temporal reasoning limitations. To mitigate shortcut learning, we propose a dissimilarity-based sampling strategy that selects semantically diverse segments of the trajectory during training. In real-world robotic manipulation tasks, VITA generalizes from a single training environment to diverse out-of-distribution tasks, environments, and embodiments, outperforming the state-of-the-art zero-shot method using autoregressive VLMs. Furthermore, we demonstrate that VITA’s zero-shot value estimates can be utilized for reward shaping in offline reinforcement learning, resulting in multi-task policies on the Meta-World benchmark that exceed the performance of those trained with the simulation’s fuzzy-logic dense rewards. Project website: https://chziakas.github.io/vita/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
Related papers
- Vision Language Models are In-Context Value LearnersYecheng Jason Ma, Joey Hejna, Chuyuan Fu, Dhruv Shah et al.ICLR 2025
- Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language ModelsShuai Zhao, Xiaohan Wang, Linchao Zhu, Yi YangICLR 2024 · 47 citations
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- Test-Time Zero-Shot Temporal Action LocalizationBenedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang et al.CVPR 2024
- Realistic Test-Time Adaptation of Vision-Language ModelsMaxime Zanella, Clément Fuchs, Christophe De Vleeschouwer, Ismail Ben AyedCVPR 2025
