PAC-Bayesian Offline Contextual Bandits With Guarantees
Otmane Sakhi, Pierre Alquier, Nicolas Chopin
摘要
This paper introduces a new principled approach for off-policy learning in contextual bandits. Unlike previous work, our approach does not derive learning principles from intractable or loose bounds. We analyse the problem through the PAC-Bayesian lens, interpreting policies as mixtures of decision rules. This allows us to propose novel generalization bounds and provide tractable algorithms to optimize them. We prove that the derived bounds are tighter than their competitors, and can be optimized directly to confidently improve upon the logging policy offline. Our approach learns policies with guarantees, uses all available data and does not require tuning additional hyperparameters on held-out sets. We demonstrate through extensive experiments the effectiveness of our approach in providing performance guarantees in practical scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 被引用 21 次
- Exponential Smoothing for Off-Policy LearningImad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna KorbaICML 2023 · 被引用 17 次
- Learning via Wasserstein-Based High Probability Generalisation BoundsPaul Viallard, Maxime Haddouche, Umut Simsekli, Benjamin GuedjNeurIPS 2023 · 被引用 16 次
- Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual BanditsQingyue Zhao, Kaixuan Ji, Heyang Zhao, Tong Zhang 等ICLR 2026 · 被引用 9 次
- Exploiting Similarities in A/B Testing with Off-Policy EstimationOtmane Sakhi, Alexandre Gilotte, David RohdeKDD 2026 · 被引用 2 次
它引用的顶会 Paper4
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova 等AAAI 2020 · 被引用 48 次
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 被引用 35 次
- BLOB: A Probabilistic Model for Recommendation that Combines Organic and Bandit SignalsOtmane Sakhi, Stephen Bonner, David Rohde, Flavian VasileKDD 2020 · 被引用 18 次
- Offline Contextual Bandits with Overparameterized ModelsDavid Brandfonbrener, William F. Whitney, Rajesh Ranganath, Joan BrunaICML 2021 · 被引用 12 次
相关 Paper
- Offline Multi-Objective Bandits: From Logged Data to Pareto-Optimal PoliciesJi Cheng, Song Lai, Shunyu Yao, Bo XueAAAI 2026 · 被引用 1 次
- Optimal Regret for Policy Optimization in Contextual BanditsOrin Levy, Yishay MansourICML 2026 · 被引用 1 次
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- PAC-Bayesian Reinforcement Learning Trains Generalizable PoliciesAbdelkrim ZITOUNI, Mehdi Hennequin, Juba Agoun, Ryan Horache 等ICML 2026 · 被引用 1 次
- An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear BanditsAndrea Tirinzoni, Matteo Pirotta, Marcello Restelli, Alessandro LazaricNeurIPS 2020 · 被引用 37 次
