VRL3: A Data-Driven Framework for Visual Deep Reinforcement Learning
Che Wang, Xufang Luo, Keith W. Ross, Dongsheng Li
Abstract
We propose VRL3, a powerful data-driven framework with a simple design for solving challenging visual deep reinforcement learning (DRL) tasks. We analyze a number of major obstacles in taking a data-driven approach, and present a suite of design principles, novel findings, and critical insights about data-driven visual DRL. Our framework has three stages: in stage 1, we leverage non-RL datasets (e.g. ImageNet) to learn task-agnostic visual representations; in stage 2, we use offline RL data (e.g. a limited number of expert demonstrations) to convert the task-agnostic representations into more powerful task-specific representations; in stage 3, we fine-tune the agent with online RL. On a set of challenging hand manipulation tasks with sparse reward and realistic visual inputs, compared to the previous SOTA, VRL3 achieves an average of 780% better sample efficiency. And on the hardest task, VRL3 is 1220% more sample efficient (2440% when using a wider encoder) and solves the task with only 10% of the computation. These significant results clearly demonstrate the great potential of data-driven deep reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bd7e55a-41cb-4d7c-8536-b256b7d4150fCited by top-tier papers26
- On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch BaselineNicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu et al.ICML 2023 · 78 citations
- Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?Xiang Li, Jinghuan Shang, Srijan Das, Michael S. RyooNeurIPS 2022 · 43 citations
- Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training StagesGuozheng Ma, Lu Li, Sen Zhang, Zixuan Liu et al.ICLR 2024 · 32 citations
- For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created EqualYingdong Hu, Renhao Wang, Li Erran Li, Yang GaoICML 2023 · 28 citations
- Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement LearningGuozheng Ma, Linrui Zhang, Haoyu Wang, Lu Li et al.NeurIPS 2023 · 24 citations
Builds on43
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Efficient Reinforcement Learning by Guiding World Models with Non-Curated DataYi Zhao, Aidan Scannell, Wenshuai Zhao, Yuxin Hou et al.ICLR 2026 · 2 citations
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu et al.ICML 2025
- MoDem: Accelerating Visual Model-Based Reinforcement Learning with DemonstrationsNicklas Hansen, Yixin Lin, Hao Su, Xiaolong Wang et al.ICLR 2023 · 8 citations
- H-InDex: Visual Reinforcement Learning with Hand-Informed Representations for Dexterous ManipulationYanjie Ze, Yuyao Liu, Ruizhe Shi, Jiaxin Qin et al.NeurIPS 2023 · 1 citation
- Efficient Reinforcement Learning Through Adaptively Pretrained Visual EncoderYuhan Zhang, Guoqing Ma, Guangfu Hao, Liangxuan Guo et al.AAAI 2025 · 3 citations
