Supervised Off-Policy Ranking
Yue Jin, Yue Zhang, Tao Qin, Xudong Zhang, Jian Yuan, Houqiang Li, Tie-Yan Liu
Abstract
Off-policy evaluation (OPE) is to evaluate a target policy with data generated by other policies. Most previous OPE methods focus on precisely estimating the true performance of a policy. We observe that in many applications, (1) the end goal of OPE is to compare two or multiple candidate policies and choose a good one, which is a much simpler task than precisely evaluating their true performance; and (2) there are usually multiple policies that have been deployed to serve users in real-world systems and thus the true performance of these policies can be known. Inspired by the two observations, in this work, we study a new problem, supervised off-policy ranking (SOPR), which aims to rank a set of target policies based on supervised learning by leveraging off-policy data and policies with known performance. We propose a method to solve SOPR, which learns a policy scoring model by minimizing a ranking loss of the training policies rather than estimating the precise policy performance. The scoring model in our method, a hierarchical Transformer based model, maps a set of state-action pairs to a score, where the state of each pair comes from the off-policy data and the action is taken by a target policy on the state in an offline manner. Extensive experiments on public datasets show that our method outperforms baseline methods in terms of rank correlation, regret value, and stability. Our code is publicly available at GitHub 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca88f879-a6a3-4eef-a6d9-039f224fb1c5Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
- Multi-Decoder Attention Model with Embedding Glimpse for Solving Vehicle Routing ProblemsLiang Xin, Wen Song, Zhiguang Cao, Jie ZhangAAAI 2021 · 209 citations
- Off-Policy Evaluation via the Regularized LagrangianMengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li et al.NeurIPS 2020 · 125 citations
- Benchmarks for Deep Off-Policy EvaluationJustin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker et al.ICLR 2021 · 112 citations
Related papers
- Policy-Adaptive Estimator Selection for Off-Policy EvaluationTakuma Udagawa, Haruka Kiyohara, Yusuke Narita, Yuta Saito et al.AAAI 2023 · 29 citations
- Off-Policy Evaluation of Ranking Policies under Diverse User BehaviorHaruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu et al.KDD 2023 · 8 citations
- Off-Policy Evaluation for Ranking Policies under Deterministic Logging PoliciesKoichi Tanaka, Kazuki Kawamura, Takanori Muroi, Yusuke Narita et al.ICLR 2026 · 1 citation
- OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple EstimatorsAllen Nie, Yash Chandak, Christina J. Yuan, Anirudhan Badrinath et al.NeurIPS 2024 · 7 citations
- Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy EvaluationHaruka Kiyohara, Ren Kishimoto, Kosuke Kawakami, Ken Kobayashi et al.ICLR 2024 · 15 citations
