Digi-Q: Learning VLM Q-Value Functions for Training Device-Control Agents
Hao Bai, Yifei Zhou, Li Erran Li, Sergey Levine, Aviral Kumar
Abstract
While a number of existing approaches for building foundation model agents rely on prompting or fine-tuning with human demonstrations, it is not sufficient in dynamic environments (e.g., mobile device control). On-policy reinforcement learning (RL) should address these limitations, but collecting actual rollouts in an environment is often undesirable in truly open-ended agentic problems such as mobile device control or interacting with humans, where each unit of interaction is associated with a cost. In such scenarios, a method for policy learning that can utilize off-policy experience by learning a trained action-value function is much more effective. In this paper, we develop an approach, called Digi-Q, to train VLM-based action-value Q-functions which are then used to extract the agent policy. We study our approach in the mobile device control setting. Digi-Q trains the Q-function using offline temporal-difference (TD) learning, on top of frozen, intermediate-layer features of a VLM. Compared to fine-tuning the whole VLM, this approach saves us compute and enhances scalability. To make the VLM features amenable for representing the Q-function, we need to employ an initial phase of fine-tuning to amplify coverage over actionable information needed for value function. Once trained, we use this Q-function via a Best-of-N policy extraction operator that imitates the best action out of multiple candidate actions from the current policy as ranked by the value function, enabling policy improvement without environment interaction. Digi-Q outperforms several prior methods on user-scale device control tasks in Android-in-the-Wild, attaining 21.2% improvement over prior best-performing method. In some cases, our Digi-Q approach already matches state-of-the-art RL methods that require interaction. The project is open-sourced at https://github.com/DigiRL-agent/digiq Recently, the community has been turning towards using reinforcement learning (RL) methods for training agentic policies. RL avoids the shortcomings of imitation and prompting, by explicitly training the policy to solve tasks (Zhou et al., 2024b; Verma et al., 2022; Snell et al., 2023; Abdulhai et al., 2023) . That said, the best performing RL methods today for improving a policy in multi-step agentic tasks rely critically on interaction due to the use of policy gradient updates (Yao et al., 2023) coupled with Monte-Carlo values (Bai et al., 2024; Putta et al., 2024; Shao et al., 2024) , which often require sufficient amounts of on-policy data to get a low-variance learning signal. The amount of on-policy data needed is likely only larger in non-stationary and dynamic environments (Bai et al., 2024) . If on the other hand, we could train a critic (i.e., an action-value function) that could score a policy's
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple ActionsGuo Gan, Yuxuan Ding, Cong Chen, Yuwei Ren et al.ACL 2026 · 6 citations
- LBM: Hierarchical Large Auto-Bidding Model via Reasoning and ActingYewen Li, Zhiyi Lyu, Peng Jiang, Qingpeng Cai et al.WWW 2026
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement LearningLang Feng, Weihao Tan, Zhiyi Lyu, Longtao Zheng et al.ICML 2025
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun et al.ICML 2024 · 496 citations
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningHao Bai, Yifei Zhou, Jiayi Pan, Mert Cemri et al.NeurIPS 2024 · 239 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
Related papers
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Offline Q-learning on Diverse Multi-Task Data Both Scales And GeneralizesAviral Kumar, Rishabh Agarwal, Xinyang Geng, George Tucker et al.ICLR 2023 · 3 citations
- RLIF: Interactive Imitation Learning as Reinforcement LearningJianlan Luo, Perry Dong, Yuexiang Zhai, Yi Ma et al.ICLR 2024 · 31 citations
- Is Value Learning Really the Main Bottleneck in Offline RL?Seohong Park, Kevin Frans, Sergey Levine, Aviral KumarNeurIPS 2024 · 99 citations
- Pretraining a Shared Q-Network for Data-Efficient Offline Reinforcement LearningJongchan Park, Mingyu Park, Donghwan LeeNeurIPS 2025 · 2 citations
