Adaptive Exploration for Data-Efficient General Value Function Evaluations
Arushi Jain, Josiah Hanna, Doina Precup
Abstract
General Value Functions (GVFs) (Sutton et al., 2011) represent predictive knowledge in reinforcement learning. Each GVF computes the expected return for a given policy, based on a unique reward. Existing methods relying on fixed behavior policies or pre-collected data often face data efficiency issues when learning multiple GVFs in parallel using off-policy methods. To address this, we introduce GVFExplorer, which adaptively learns a single behavior policy that efficiently collects data for evaluating multiple GVFs in parallel. Our method optimizes the behavior policy by minimizing the total variance in return across GVFs, thereby reducing the required environmental interactions. We use an existing temporal-difference-style variance estimator to approximate the return variance. We prove that each behavior policy update decreases the overall mean squared error in GVF predictions. We empirically show our method's performance in tabular and nonlinear function approximation settings, including Mujoco environments, with stationary and non-stationary reward signals, optimizing data usage and reducing prediction errors across multiple GVFs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f130c57d-37af-4a26-a910-de3e109bb9ccCited by top-tier papers1
Ask how each one uses itBuilds on4
- BYOL-Explore: Exploration by Bootstrapped PredictionZhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires et al.NeurIPS 2022 · 104 citations
- Continual Auxiliary Task LearningMatthew McLeod, Chunlok Lo, Matthew Schlegel, Andrew Jacobsen et al.NeurIPS 2021 · 13 citations
- Variance Penalized On-Policy and Off-Policy Actor-CriticArushi Jain, Gandharv Patil, Ayush Jain, Khimya Khetarpal et al.AAAI 2021 · 11 citations
- A Unifying Framework of Off-Policy General Value Function EvaluationTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangNeurIPS 2022 · 2 citations
Related papers
- Gamma-Nets: Generalizing Value Estimation over TimescaleCraig Sherstan, Shibhansh Dohare, James MacGlashan, Johannes Günther et al.AAAI 2020 · 14 citations
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 29 citations
- Efficient Multi-Policy Evaluation for Reinforcement LearningShuze Daniel Liu, Claire Chen, Shangtong ZhangAAAI 2025 · 2 citations
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 7 citations
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 56 citations
