Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement Learning
Rujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht, Josiah Hanna
摘要
Reinforcement learning (RL) algorithms are often categorized as either on-policy or off-policy depending on whether they use data from a target policy of interest or from a different behavior policy. In this paper, we study a subtle distinction between on-policy data and on-policy sampling in the context of the RL sub-problem of policy evaluation. We observe that on-policy sampling may fail to match the expected distribution of on-policy data after observing only a finite number of trajectories and this failure hinders data-efficient policy evaluation. Towards improved data-efficiency, we show how non-i.i.d., off-policy sampling can produce data that more closely matches the expected on-policy data distribution and consequently increases the accuracy of the Monte Carlo estimator for policy evaluation. We introduce a method called Robust On-Policy Sampling and demonstrate theoretically and empirically that it produces data that converges faster to the expected on-policy distribution compared to on-policy sampling. Empirically, we show that this faster convergence leads to lower mean squared error policy value estimates.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Conditional Mutual Information for Disentangled Representations in Reinforcement LearningMhairi Dunion, Trevor McInroe, Kevin Sebastian Luck, Josiah Hanna 等NeurIPS 2023 · 被引用 41 次
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 被引用 7 次
- Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive ApproachRiccardo Poiani, Nicole Nobili, Alberto Maria Metelli, Marcello RestelliNeurIPS 2023 · 被引用 3 次
- Off-Policy Selection for Initiating Human-Centric Experimental DesignGe Gao, Xi Yang, Qitong Gao, Song Ju 等NeurIPS 2024 · 被引用 1 次
- Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement LearningAlexander W. Goodall, Edwin Hamel-De le Court, Francesco BelardinelliAAAI 2026 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Imitation Learning via Off-Policy Distribution MatchingIlya Kostrikov, Ofir Nachum, Jonathan TompsonICLR 2020 · 被引用 239 次
- Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-offZichen Vincent Zhang, Johannes Kirschner, Junxi Zhang, Francesco Zanini 等NeurIPS 2023 · 被引用 3 次
- Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic PoliciesNathan Kallus, Masatoshi UeharaNeurIPS 2020 · 被引用 16 次
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 被引用 60 次
