Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy Learning
Aurélien Bibaut, Nathan Kallus, Maria Dimakopoulou, Antoine Chambaz, Mark J. van der Laan
Abstract
Empirical risk minimization (ERM) is the workhorse of machine learning, whether for classification and regression or for off-policy policy learning, but its modelagnostic guarantees can fail when we use adaptively collected data, such as the result of running a contextual bandit algorithm. We study a generic importance sampling weighted ERM algorithm for using adaptively collected data to minimize the average of a loss function over a hypothesis class and provide first-of-their-kind generalization guarantees and fast convergence rates. Our results are based on a new maximal inequality that carefully leverages the importance sampling structure to obtain rates with the right dependence on the exploration rate in the data. For regression, we provide fast rates that leverage the strong convexity of squared-error loss. For policy learning, we provide rate-optimal regret guarantees that close an open gap in the existing literature whenever exploration decays to zero, as is the case for bandit-collected data. An empirical investigation validates our theory. * Alphabetical order Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdd20714-b8d8-481b-b330-4762f5af449dCited by top-tier papers4
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz et al.NeurIPS 2021 · 58 citations
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 36 citations
- Sequential Counterfactual Risk MinimizationHoussam Zenati, Eustache Diemert, Matthieu Martin, Julien Mairal et al.ICML 2023 · 6 citations
- Distributionally Robust Policy Learning under Concept DriftsJingyuan Wang, Zhimei Ren, Ruohan Zhan, Zhengyuan ZhouICML 2025
Builds on2
Related papers
- Contextual Linear Optimization with Bandit FeedbackYichun Hu, Nathan Kallus, Xiaojie Mao, Yanchen WuNeurIPS 2024
- Logarithmic Smoothing for Pessimistic Off-Policy Evaluation, Selection and LearningOtmane Sakhi, Imad Aouali, Pierre Alquier, Nicolas ChopinNeurIPS 2024 · 21 citations
- Learning from Biased Data: A Semi-Parametric ApproachPatrice Bertail, Stéphan Clémençon, Yannick Guyonvarch, Nathan NoiryICML 2021 · 6 citations
- Importance Weighted Actor-Critic for Optimal Conservative Offline Reinforcement LearningHanlin Zhu, Paria Rashidinejad, Jiantao JiaoNeurIPS 2023 · 21 citations
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
