Model Selection in Batch Policy Optimization
Jonathan Lee, George Tucker, Ofir Nachum, Bo Dai
摘要
We study the problem of model selection in batch policy optimization: given a fixed, partial-feedback dataset and model classes, learn a policy with performance that is competitive with the policy derived from the best model class. We formalize the problem in the contextual bandit setting with linear model classes by identifying three sources of error that any model selection algorithm should optimally trade-off in order to be competitive: (1) approximation error, (2) statistical complexity, and (3) coverage. The first two sources are common in model selection for supervised learning, where optimally trading-off these properties is well-studied. In contrast, the third source is unique to batch policy optimization and is due to dataset shift inherent to the setting. We first show that no batch policy optimization algorithm can achieve a guarantee addressing all three simultaneously, revealing a stark contrast between difficulties in batch policy optimization and the positive results available in supervised learning. Despite this negative result, we show that relaxing any one of the three error sources enables the design of algorithms achieving near-oracle inequalities for the remaining two. We conclude with experiments demonstrating the efficacy of these algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Data-Efficient Pipeline for Offline Reinforcement Learning with Limited DataAllen Nie, Yannis Flet-Berliac, Deon R. Jordan, William Steenbergen 等NeurIPS 2022 · 被引用 18 次
- Bellman Residual Orthogonalization for Offline Reinforcement LearningAndrea Zanette, Martin J. WainwrightNeurIPS 2022 · 被引用 14 次
- Oracle Inequalities for Model Selection in Offline Reinforcement LearningJonathan N. Lee, George Tucker, Ofir Nachum, Bo Dai 等NeurIPS 2022 · 被引用 14 次
- Model Selection for Off-policy Evaluation: New Algorithms and Experimental ProtocolPai Liu, Lingfeng Zhao, Shivangi Agarwal, Jinghan Liu 等NeurIPS 2025 · 被引用 6 次
- Cross-Validated Off-Policy EvaluationMatej Cief, Branislav Kveton, Michal KompanAAAI 2025 · 被引用 2 次
它引用的顶会 Paper12
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 等NeurIPS 2021 · 被引用 339 次
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 被引用 241 次
- Minimax-Optimal Off-Policy Evaluation with Linear Function ApproximationYaqi Duan, Zeyu Jia, Mengdi WangICML 2020 · 被引用 161 次
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement LearningAndrea Zanette, Martin J. Wainwright, Emma BrunskillNeurIPS 2021 · 被引用 140 次
相关 Paper
- Almost Optimal Batch-Regret Tradeoff for Batch Linear Contextual BanditsZihan Zhang, Xiangyang Ji, Yuan ZhouICLR 2025
- On the Optimality of Batch Policy Optimization AlgorithmsChenjun Xiao, Yifan Wu, Jincheng Mei, Bo Dai 等ICML 2021 · 被引用 36 次
- The Pareto Frontier of model selection for general Contextual BanditsTeodor Vanislavov Marinov, Julian ZimmertNeurIPS 2021 · 被引用 31 次
- Transportability for Bandits with Data from Different EnvironmentsAlexis Bellot, Alan Malek, Silvia ChiappaNeurIPS 2023 · 被引用 11 次
- Regret Minimization with Performative FeedbackMeena Jagadeesan, Tijana Zrnic, Celestine Mendler-DünnerICML 2022 · 被引用 41 次
