VIPeR: Provably Efficient Algorithm for Offline RL with Neural Function Approximation
Thanh Nguyen-Tang, Raman Arora
Abstract
We propose a novel algorithm for offline reinforcement learning called Value Iteration with Perturbed Rewards (VIPeR), which amalgamates the pessimism principle with random perturbations of the value function. Most current offline RL algorithms explicitly construct statistical confidence regions to obtain pessimism via lower confidence bounds (LCB), which cannot easily scale to complex problems where a neural network is used to estimate the value functions. Instead, VIPeR implicitly obtains pessimism by simply perturbing the offline data multiple times with carefully-designed i.i.d. Gaussian noises to learn an ensemble of estimated state-action value functions and acting greedily with respect to the minimum of the ensemble. The estimated state-action values are obtained by fitting a parametric model (e.g., neural networks) to the perturbed datasets using gradient descent. As a result, VIPeR only needs time complexity for action selection, while LCB-based algorithms require at least , where is the total number of trajectories in the offline data. We also propose a novel data-splitting technique that helps remove a factor involving the log of the covering number in our bound. We prove that VIPeR yields a provable uncertainty quantifier with overparameterized neural networks and enjoys a bound on sub-optimality of , where is the effective dimension, is the horizon length and measures the distributional shift. We corroborate the statistical and computational efficiency of VIPeR with an empirical evaluation on a wide set of synthetic and real-world datasets. To the best of our knowledge, VIPeR is the first algorithm for offline RL that is provably efficient for general Markov decision processes (MDPs) with neural network function approximation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94e5542d-0728-4c32-bbdc-1629e1d3c7c5Cited by top-tier papers4
- Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement LearningQiwei Di, Heyang Zhao, Jiafan He, Quanquan GuICLR 2024 · 9 citations
- On Sample-Efficient Offline Reinforcement Learning: Data Diversity, Posterior Sampling and BeyondThanh Nguyen-Tang, Raman AroraNeurIPS 2023 · 7 citations
- Online Optimization for Offline Safe Reinforcement LearningYassine Chemingui, Aryan Deshwal, Alan Fern, Thanh Nguyen-Tang et al.NeurIPS 2025 · 3 citations
- Information-Directed Pessimism for Offline Reinforcement LearningAlec Koppel, Sujay Bhatt, Jiacheng Guo, Joe Eappen et al.ICML 2024 · 3 citations
Builds on32
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 430 citations
Related papers
- Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision ProcessesMiao Lu, Yifei Min, Zhaoran Wang, Zhuoran YangICLR 2023
- Offline Reinforcement Learning with Differentiable Function Approximation is Provably EfficientMing Yin, Mengdi Wang, Yu-Xiang WangICLR 2023
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- VIPO: Value Function Inconsistency Penalized Offline Reinforcement LearningXuyang Chen, Keyu Yan, Guojian Wang, Lin ZhaoICML 2026 · 3 citations
- Bi-Level Offline Policy Optimization with Limited ExplorationWenzhuo ZhouNeurIPS 2023 · 6 citations
