Thompson Sampling Efficiently Learns to Control Diffusion Processes
Mohamad Kazem Shirani Faradonbeh, Mohamad Sadegh Shirani Faradonbeh, Mohsen Bayati
Abstract
Linear diffusion processes serve as canonical continuous-time models for dynamic decision-making under uncertainty. These systems evolve according to drift matrices that specify the instantaneous rates of change in the expected system state, while also experiencing continuous random disturbances modeled by Brownian noise. For instance, in medical applications such as artificial pancreas systems, the drift matrices represent the internal dynamics of glucose concentrations and how insulin levels influence them. Classical results in stochastic control provide optimal policies under perfect knowledge of the drift matrices. However, practical decision-making scenarios typically feature uncertainty about the drift; in medical contexts, such parameters are patient-specific and unknown, requiring adaptive policies capable of efficiently learning the drift matrices while simultaneously ensuring system stability and optimal performance. We study the popular Thompson sampling algorithm for decision-making in linear diffusion processes with unknown drift matrices. For this algorithm that designs control policies as if samples from a posterior belief about the parameters fully coincide with the unknown truth, we establish efficiency. That is, Thompson sampling learns optimal control actions fast, incurring only a square-root of time regret, and also learns to stabilize the system in a short time period. To our knowledge, this is the first such result for Thompson sampling in a diffusion process control problem. Moreover, our empirical simulations in three settings that involve blood-glucose and flight control demonstrate that Thompson sampling significantly improves regret, compared to the state-of-the-art algorithms, suggesting it explores in a more guarded fashion. Our theoretical analysis includes characterization of a certain optimality manifold that relates the geometry of the drift matrices to the optimal control of the diffusion process, among others. We expect the technical contributions to be of independent interest in the study of decision-making under uncertainty problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear PredictionWeiyi Lv, Yuhang Huang, Ning Zhang, Ruei-Sung Lin et al.CVPR 2024 · 36 citations
- Online Control with Adversarial Disturbance for Continuous-time Linear SystemsJingwei Li, Jing Dong, Can Chang, Baoxiang Wang et al.NeurIPS 2024 · 1 citation
- Finite Sample Analyses for Continuous-time Linear Systems: System Identification and Online ControlHongyi Zhou, Jingwei Li, Jingzhao ZhangNeurIPS 2025
Related papers
- Causal Modeling of Policy Interventions From Treatment-Outcome SequencesCaglar Hizli, S. T. John, Anne Tuulikki Juuti, Tuure Tapani Saarinen et al.ICML 2023 · 7 citations
- Bayesian Learning of Optimal Policies in Markov Decision Processes with Countably Infinite State-SpaceSaghar Adler, Vijay G. SubramanianNeurIPS 2023 · 4 citations
- Quasi-optimal Reinforcement Learning with Continuous ActionsYuhan Li, Wenzhuo Zhou, Ruoqing ZhuICLR 2023 · 2 citations
- ESCADA: Efficient Safety and Context Aware Dose Allocation for Precision MedicineIlker Demirel, Ahmet Alparslan Celik, Cem TekinNeurIPS 2022 · 6 citations
- Thompson Sampling with Diffusion Generative PriorYu-Guan Hsieh, Shiva Prasad Kasiviswanathan, Branislav Kveton, Patrick BlöbaumICML 2023 · 7 citations
