Thompson Sampling Efficiently Learns to Control Diffusion Processes
Mohamad Kazem Shirani Faradonbeh, Mohamad Sadegh Shirani Faradonbeh, Mohsen Bayati
摘要
Linear diffusion processes serve as canonical continuous-time models for dynamic decision-making under uncertainty. These systems evolve according to drift matrices that specify the instantaneous rates of change in the expected system state, while also experiencing continuous random disturbances modeled by Brownian noise. For instance, in medical applications such as artificial pancreas systems, the drift matrices represent the internal dynamics of glucose concentrations and how insulin levels influence them. Classical results in stochastic control provide optimal policies under perfect knowledge of the drift matrices. However, practical decision-making scenarios typically feature uncertainty about the drift; in medical contexts, such parameters are patient-specific and unknown, requiring adaptive policies capable of efficiently learning the drift matrices while simultaneously ensuring system stability and optimal performance. We study the popular Thompson sampling algorithm for decision-making in linear diffusion processes with unknown drift matrices. For this algorithm that designs control policies as if samples from a posterior belief about the parameters fully coincide with the unknown truth, we establish efficiency. That is, Thompson sampling learns optimal control actions fast, incurring only a square-root of time regret, and also learns to stabilize the system in a short time period. To our knowledge, this is the first such result for Thompson sampling in a diffusion process control problem. Moreover, our empirical simulations in three settings that involve blood-glucose and flight control demonstrate that Thompson sampling significantly improves regret, compared to the state-of-the-art algorithms, suggesting it explores in a more guarded fashion. Our theoretical analysis includes characterization of a certain optimality manifold that relates the geometry of the drift matrices to the optimal control of the diffusion process, among others. We expect the technical contributions to be of independent interest in the study of decision-making under uncertainty problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear PredictionWeiyi Lv, Yuhang Huang, Ning Zhang, Ruei-Sung Lin 等CVPR 2024 · 被引用 36 次
- Online Control with Adversarial Disturbance for Continuous-time Linear SystemsJingwei Li, Jing Dong, Can Chang, Baoxiang Wang 等NeurIPS 2024 · 被引用 1 次
- Finite Sample Analyses for Continuous-time Linear Systems: System Identification and Online ControlHongyi Zhou, Jingwei Li, Jingzhao ZhangNeurIPS 2025
相关 Paper
- Causal Modeling of Policy Interventions From Treatment-Outcome SequencesCaglar Hizli, S. T. John, Anne Tuulikki Juuti, Tuure Tapani Saarinen 等ICML 2023 · 被引用 7 次
- Bayesian Learning of Optimal Policies in Markov Decision Processes with Countably Infinite State-SpaceSaghar Adler, Vijay G. SubramanianNeurIPS 2023 · 被引用 4 次
- Quasi-optimal Reinforcement Learning with Continuous ActionsYuhan Li, Wenzhuo Zhou, Ruoqing ZhuICLR 2023 · 被引用 2 次
- ESCADA: Efficient Safety and Context Aware Dose Allocation for Precision MedicineIlker Demirel, Ahmet Alparslan Celik, Cem TekinNeurIPS 2022 · 被引用 6 次
- Thompson Sampling with Diffusion Generative PriorYu-Guan Hsieh, Shiva Prasad Kasiviswanathan, Branislav Kveton, Patrick BlöbaumICML 2023 · 被引用 7 次
