Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement Learning
Rujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht, Josiah Hanna
Abstract
Reinforcement learning (RL) algorithms are often categorized as either on-policy or off-policy depending on whether they use data from a target policy of interest or from a different behavior policy. In this paper, we study a subtle distinction between on-policy data and on-policy sampling in the context of the RL sub-problem of policy evaluation. We observe that on-policy sampling may fail to match the expected distribution of on-policy data after observing only a finite number of trajectories and this failure hinders data-efficient policy evaluation. Towards improved data-efficiency, we show how non-i.i.d., off-policy sampling can produce data that more closely matches the expected on-policy data distribution and consequently increases the accuracy of the Monte Carlo estimator for policy evaluation. We introduce a method called Robust On-Policy Sampling and demonstrate theoretically and empirically that it produces data that converges faster to the expected on-policy distribution compared to on-policy sampling. Empirically, we show that this faster convergence leads to lower mean squared error policy value estimates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Conditional Mutual Information for Disentangled Representations in Reinforcement LearningMhairi Dunion, Trevor McInroe, Kevin Sebastian Luck, Josiah Hanna et al.NeurIPS 2023 · 41 citations
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 7 citations
- Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive ApproachRiccardo Poiani, Nicole Nobili, Alberto Maria Metelli, Marcello RestelliNeurIPS 2023 · 3 citations
- Off-Policy Selection for Initiating Human-Centric Experimental DesignGe Gao, Xi Yang, Qitong Gao, Song Ju et al.NeurIPS 2024 · 1 citation
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 1 citation
Builds on2
Related papers
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 239 citations
- Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-offZichen Vincent Zhang, Johannes Kirschner, Junxi Zhang, Francesco Zanini et al.NeurIPS 2023 · 3 citations
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 16 citations
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 60 citations
