Safe Exploration for Efficient Policy Evaluation and Comparison
Runzhe Wan, Branislav Kveton, Rui Song
Abstract
High-quality data plays a central role in ensuring the accuracy of policy evaluation. This paper initiates the study of efficient and safe data collection for bandit policy evaluation. We formulate the problem and investigate its several representative variants. For each variant, we analyze its statistical properties, derive the corresponding exploration policy, and design an efficient algorithm for computing it. Both theoretical analysis and experiments support the usefulness of the proposed methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36c4f733-84a4-4462-95ca-6b1391e4465eCited by top-tier papers8
- Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision MakingTing Li, Chengchun Shi, Jianing Wang, Fan Zhou et al.NeurIPS 2023 · 21 citations
- Experiment Planning with Function ApproximationAldo Pacchiano, Jonathan Lee, Emma BrunskillNeurIPS 2023 · 6 citations
- Exploiting Similarities in A/B Testing with Off-Policy EstimationOtmane Sakhi, Alexandre Gilotte, David RohdeKDD 2026 · 2 citations
- A More Accurate Algorithm Comparison through A/B Testing using Offline Evaluation MethodsKoki Konishi, Masataka Ushiku, Yuta SaitoKDD 2026 · 1 citation
- Designing Time Series Experiments in A/B Testing with Transformer Reinforcement LearningXiangkun Wu, Qianglin Wen, Yingying Zhang, Hongtu Zhu et al.ICLR 2026 · 1 citation
Builds on11
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Meta-Thompson SamplingBranislav Kveton, Mikhail Konobeev, Manzil Zaheer, Chih-Wei Hsu et al.ICML 2021 · 74 citations
- Optimal Off-Policy Evaluation from Multiple Logging PoliciesNathan Kallus, Yuta Saito, Masatoshi UeharaICML 2021 · 44 citations
- Deeply-Debiased Off-Policy Interval EstimationChengchun Shi, Runzhe Wan, Victor Chernozhukov, Rui SongICML 2021 · 43 citations
- Stage-wise Conservative Linear BanditsAhmadreza Moradipari, Christos Thrampoulidis, Mahnoosh AlizadehNeurIPS 2020 · 37 citations
Related papers
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
- SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDPSubhojyoti Mukherjee, Josiah P. Hanna, Robert D. NowakICML 2024
- Efficient Policy Evaluation with Safety Constraint for Reinforcement LearningClaire Chen, Shuze Daniel Liu, Shangtong ZhangICLR 2025
- Doubly Optimal Policy Evaluation for Reinforcement LearningShuze Daniel Liu, Claire Chen, Shangtong ZhangICLR 2025
- Efficient Policy Evaluation with Offline Data Informed Behavior Policy DesignShuze Daniel Liu, Shangtong ZhangICML 2024 · 7 citations
