PAC-Bayesian Offline Contextual Bandits With Guarantees
Otmane Sakhi, Pierre Alquier, Nicolas Chopin
Abstract
This paper introduces a new principled approach for off-policy learning in contextual bandits. Unlike previous work, our approach does not derive learning principles from intractable or loose bounds. We analyse the problem through the PAC-Bayesian lens, interpreting policies as mixtures of decision rules. This allows us to propose novel generalization bounds and provide tractable algorithms to optimize them. We prove that the derived bounds are tighter than their competitors, and can be optimized directly to confidently improve upon the logging policy offline. Our approach learns policies with guarantees, uses all available data and does not require tuning additional hyperparameters on held-out sets. We demonstrate through extensive experiments the effectiveness of our approach in providing performance guarantees in practical scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Exponential Smoothing for Off-Policy LearningImad Aouali, Victor-Emmanuel Brunel, David Rohde, Anna KorbaICML 2023 · 17 citations
- Learning via Wasserstein-Based High Probability Generalisation BoundsPaul Viallard, Maxime Haddouche, Umut Simsekli, Benjamin GuedjNeurIPS 2023 · 16 citations
- Towards a Sharp Analysis of Offline Policy Learning for -Divergence-Regularized Contextual BanditsQingyue Zhao, Kaixuan Ji, Heyang Zhao, Tong Zhang et al.ICLR 2026 · 9 citations
- Exploiting Similarities in A/B Testing with Off-Policy EstimationOtmane Sakhi, Alexandre Gilotte, David RohdeKDD 2026 · 2 citations
Builds on4
- Distributionally Robust Counterfactual Risk MinimizationLouis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova et al.AAAI 2020 · 48 citations
- Offline Neural Contextual Bandits: Pessimism, Optimization and GeneralizationThanh Nguyen-Tang, Sunil Gupta, A. Tuan Nguyen, Svetha VenkateshICLR 2022 · 35 citations
- BLOB: A Probabilistic Model for Recommendation that Combines Organic and Bandit SignalsOtmane Sakhi, Stephen Bonner, David Rohde, Flavian VasileKDD 2020 · 18 citations
- Offline Contextual Bandits with Overparameterized ModelsDavid Brandfonbrener, William F. Whitney, Rajesh Ranganath, Joan BrunaICML 2021 · 12 citations
Related papers
- Offline Multi-Objective Bandits: From Logged Data to Pareto-Optimal PoliciesJi Cheng, Song Lai, Shunyu Yao, Bo XueAAAI 2026 · 1 citation
- Optimal Regret for Policy Optimization in Contextual BanditsOrin Levy, Yishay MansourICML 2026 · 1 citation
- Off-Policy Learning in Large Action Spaces: Optimization Matters More Than EstimationImad AOUALI, Otmane SakhiICML 2026
- PAC-Bayesian Reinforcement Learning Trains Generalizable PoliciesAbdelkrim ZITOUNI, Mehdi Hennequin, Juba Agoun, Ryan Horache et al.ICML 2026 · 1 citation
- An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear BanditsAndrea Tirinzoni, Matteo Pirotta, Marcello Restelli, Alessandro LazaricNeurIPS 2020 · 37 citations
