Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation
Haruka Kiyohara, Ren Kishimoto, Kosuke Kawakami, Ken Kobayashi, Kazuhide Nakata, Yuta Saito
Abstract
Off-Policy Evaluation (OPE) aims to assess the effectiveness of counterfactual policies using only offline logged data and is often used to identify the top-k promising policies for deployment in online A/B tests. Existing evaluation metrics for OPE estimators primarily focus on the"accuracy"of OPE or that of downstream policy selection, neglecting risk-return tradeoff in the subsequent online policy deployment. To address this issue, we draw inspiration from portfolio evaluation in finance and develop a new metric, called SharpeRatio@k, which measures the risk-return tradeoff of policy portfolios formed by an OPE estimator under varying online evaluation budgets (k). We validate our metric in two example scenarios, demonstrating its ability to effectively distinguish between low-risk and high-risk estimators and to accurately identify the most efficient one. Efficiency of an estimator is characterized by its capability to form the most advantageous policy portfolios, maximizing returns while minimizing risks during online deployment, a nuance that existing metrics typically overlook. To facilitate a quick, accurate, and consistent evaluation of OPE via SharpeRatio@k, we have also integrated this metric into an open-source software, SCOPE-RL (https://github.com/hakuhodo-technologies/scope-rl). Employing SharpeRatio@k and SCOPE-RL, we conduct comprehensive benchmarking experiments on various estimators and RL tasks, focusing on their risk-return tradeoff. These experiments offer several interesting directions and suggestions for future OPE research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 187278b0-144d-4d8b-8e6d-edb3376a589fCited by top-tier papers7
- Off-Policy Evaluation of Slate Bandit Policies via Optimizing AbstractionHaruka Kiyohara, Masahiro Nomura, Yuta SaitoWWW 2024 · 18 citations
- Long-term Off-Policy Evaluation and LearningYuta Saito, Himan Abdollahpouri, Jesse Anderton, Ben Carterette et al.WWW 2024 · 15 citations
- Off-Policy Learning with Limited SupplyKoichi Tanaka, Ren Kishimoto, Bushun Kawagishi, Yusuke Narita et al.WWW 2026
- POTEC: Off-Policy Contextual Bandits for Large Action Spaces via Policy DecompositionYuta Saito, Jihan Yao, Thorsten JoachimsICLR 2025
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng et al.ISSTA 2026
Builds on14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Minimax Weight and Q-Function Learning for Off-Policy EvaluationMasatoshi Uehara, Jiawei Huang, Nan JiangICML 2020 · 199 citations
- Deployment-Efficient Reinforcement Learning via Model-Based Offline OptimizationTatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum et al.ICLR 2021 · 166 citations
- Batch Value-function Approximation with Only RealizabilityTengyang Xie, Nan JiangICML 2021 · 131 citations
Related papers
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath et al.NeurIPS 2024 · 7 citations
- Supervised Off-Policy RankingYue Jin, Yue Zhang, Tao Qin, Xudong Zhang et al.ICML 2022 · 6 citations
- State-Action Similarity-Based Representations for Off-Policy EvaluationBrahma S. Pavse, Josiah HannaNeurIPS 2023 · 5 citations
- Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftMasatoshi Uehara, Masahiro Kato, Shota YasuiNeurIPS 2020 · 60 citations
