Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
Imad AOUALI, Otmane Sakhi
摘要
Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimators with improved statistical properties, assuming that better estimators inherently yield superior policies. Although theoretically justified, this estimator-centric approach neglects a critical practical obstacle: challenging optimization landscapes. In this paper, we provide theoretical insights and empirical evidence showing that current OPL methods encounter severe optimization issues, particularly as the action space grows. We show that estimator-aware policy parametrization can mitigate, but not fully resolve, optimization challenges. Building on this, we explore simpler weighted log-likelihood objectives and demonstrate that they enjoy substantially better optimization properties and still recover competitive, often superior, learned policies. Our findings emphasize the necessity of explicitly addressing optimization considerations in the development of OPL algorithms for large action spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 被引用 128 次
- Adapting to Misspecification in Contextual BanditsDylan J. Foster, Claudio Gentile, Mehryar Mohri, Julian ZimmertNeurIPS 2020 · 被引用 111 次
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 被引用 62 次
相关 Paper
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- PAC-Bayesian Offline Contextual Bandits With GuaranteesOtmane Sakhi, Pierre Alquier, Nicolas ChopinICML 2023 · 被引用 23 次
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 被引用 55 次
- Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsHaanvid Lee, Jongmin Lee, Yunseon Choi, Wonseok Jeon 等NeurIPS 2022 · 被引用 7 次
