Off-Policy Learning in Large Action Spaces: Optimization Matters More Than Estimation
Imad AOUALI, Otmane Sakhi
Abstract
Off-policy evaluation (OPE) and off-policy learning (OPL) are foundational for decision-making in offline contextual bandits. Recent advances in OPL primarily optimize OPE estimators with improved statistical properties, assuming that better estimators inherently yield superior policies. Although theoretically justified, this estimator-centric approach neglects a critical practical obstacle: challenging optimization landscapes. In this paper, we provide theoretical insights and empirical evidence showing that current OPL methods encounter severe optimization issues, particularly as the action space grows. We show that estimator-aware policy parametrization can mitigate, but not fully resolve, optimization challenges. Building on this, we explore simpler weighted log-likelihood objectives and demonstrate that they enjoy substantially better optimization properties and still recover competitive, often superior, learned policies. Our findings emphasize the necessity of explicitly addressing optimization considerations in the development of OPL algorithms for large action spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b58f5425-e91b-4f15-b1f2-46587830d0bbBuilds on20
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Doubly robust off-policy evaluation with shrinkageYi Su, Maria Dimakopoulou, Akshay Krishnamurthy, Miroslav DudíkICML 2020 · 128 citations
- Adapting to Misspecification in Contextual BanditsDylan J. Foster, Claudio Gentile, Mehryar Mohri, Julian ZimmertNeurIPS 2020 · 111 citations
- Off-Policy Evaluation for Large Action Spaces via EmbeddingsYuta Saito, Thorsten JoachimsICML 2022 · 62 citations
Related papers
- Cross-Domain Off-Policy Evaluation and Learning for Contextual BanditsYuta Natsubori, Masataka Ushiku, Yuta SaitoICLR 2025
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
- PAC-Bayesian Offline Contextual Bandits With GuaranteesOtmane Sakhi, Pierre Alquier, Nicolas ChopinICML 2023 · 23 citations
- Adaptive Estimator Selection for Off-Policy EvaluationYi Su, Pavithra Srinath, Akshay KrishnamurthyICML 2020 · 55 citations
- Local Metric Learning for Off-Policy Evaluation in Contextual Bandits with Continuous ActionsHaanvid Lee, Jongmin Lee, Yunseon Choi, Wonseok Jeon et al.NeurIPS 2022 · 7 citations
