The Pareto-optimal Trade-off between Regret and Statistical Inference in Linear Stochastic Bandits under Safety Constraints
Yuming Shao, Zhixuan Fang
摘要
Linear bandits traditionally prioritize regret minimization, often overlooking statistical inference of the underlying parameter as a critical objective. In high-stakes settings such as healthcare, precise parameter estimation is indispensable, as it provides fundamental insights into system mechanisms and ensures robust decision-making under covariate shift. We investigate the tripartite balance between regret, inference, and safety, deriving a fundamental minimax lower bound that characterizes the Pareto-optimal frontier of these competing goals. We then propose SERMiSC, a novel algorithm that achieves the optimal trade-off by matching this lower bound while maintaining a near-constant Õ(1) safety risk. Empirical results demonstrate that SERMiSC effectively navigates the Pareto frontier and outperforms various baselines, thereby validating our theoretical analysis.
In online learning scenarios such as clinical trials, minimiz-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Gamification of Pure Exploration for Linear BanditsRémy Degenne, Pierre Ménard, Xuedong Shang, Michal ValkoICML 2020 · 被引用 86 次
- High-Dimensional Sparse Linear BanditsBotao Hao, Tor Lattimore, Mengdi WangNeurIPS 2020 · 被引用 77 次
- An Efficient Pessimistic-Optimistic Algorithm for Stochastic Linear Bandits with General ConstraintsXin Liu, Bin Li, Pengyi Shi, Lei YingNeurIPS 2021 · 被引用 63 次
- Safe Linear Stochastic BanditsKia Khezeli, Eilyan BitarAAAI 2020 · 被引用 31 次
- Online Experimental Design With Estimation-Regret Trade-off Under Network InterferenceZhiheng Zhang, Zichen WangNeurIPS 2025 · 被引用 12 次
相关 Paper
- Design-Based Bandits Under Network Interference: Trade-Off Between Regret and Statistical InferenceZichen Wang, Haoyang Hong, Chuanhao Li, Haoxuan Li 等NeurIPS 2025 · 被引用 3 次
- Proportional Response: Contextual Bandits for Simple and Cumulative Regret MinimizationSanath Kumar Krishnamurthy, Ruohan Zhan, Susan Athey, Emma BrunskillNeurIPS 2023 · 被引用 15 次
- Privacy Preserving Adaptive Experiment DesignJiachun Li, Kaining Shi, David Simchi-LeviICML 2024 · 被引用 1 次
- Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and RiskTianrui Chen, Aditya Gangrade, Venkatesh SaligramaICML 2022 · 被引用 18 次
- Thompson Sampling for Multi-Objective Linear Contextual BanditSomangchan Park, Heesang Ann, Min-hwan OhNeurIPS 2025 · 被引用 1 次
