Maximizing the Value of Predictions in Control: Accuracy Is Not Enough
Yiheng Lin, Christopher Yeh, Zaiwei Chen, Adam Wierman
摘要
We study the value of stochastic predictions in online optimal control with random disturbances. Prior work provides performance guarantees based on prediction error but ignores the stochastic dependence between predictions and disturbances. We introduce a general framework modeling their joint distribution and define "prediction power" as the control cost improvement from the optimal use of predictions compared to ignoring the predictions. In the time-varying Linear Quadratic Regulator (LQR) setting, we derive a closed-form expression for prediction power and discuss its mismatch with prediction accuracy and connection with online policy optimization. To extend beyond LQR, we study general dynamics and costs. We establish a lower bound on prediction power under two sufficient conditions that generalize the properties of the LQR setting, characterizing the fundamental benefit of incorporating stochastic predictions. We apply this lower bound to nonquadratic costs and show that even weakly dependent predictions yield significant performance gains.
Consider a fixed predictor parameter θ. For each time step t, let I t (θ) := (W 0:t-1 , V 0:t (θ)) denote the history of past disturbances and predictions, and let F t (θ) := σ(I t (θ)) 1 . A predictive policy that applies to the predictor with parameter θ is a sequence of functions π 0:T -1 , where π t maps a state/history pair to a control action.
Given a fixed predictive policy sequence π = π 0:T -1 for a predictor parameter θ, we evaluate its performance via the expected total cost over Ξ:
), for t = 0, . . . , T -1. The optimal cost under θ is defined as J * (θ) = min π J π (θ), where the minimum is over all predictive policies that use the predictor parameter θ.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- The Power of Predictions in Online ControlChenkai Yu, Guanya Shi, Soon-Jo Chung, Yisong Yue 等NeurIPS 2020 · 被引用 88 次
- Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying SystemsYiheng Lin, Yang Hu, Guanya Shi, Haoyuan Sun 等NeurIPS 2021 · 被引用 55 次
- SODA: An Adaptive Bitrate Controller for Consistent High-Quality Video StreamingTianyu Chen, Yiheng Lin, Nicolas Christianson, Zahaib Akhtar 等SIGCOMM 2024 · 被引用 33 次
- Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and NonlinearityYiheng Lin, Yang Hu, Guannan Qu, Tongxin Li 等NeurIPS 2022 · 被引用 31 次
- Online Adaptive Policy Selection in Time-Varying Systems: No-Regret via Contractive PerturbationsYiheng Lin, James A. Preiss, Emile Anand, Yingying Li 等NeurIPS 2023 · 被引用 31 次
相关 Paper
- Optimal Dynamic Regret in LQR ControlDheeraj Baby, Yu-Xiang WangNeurIPS 2022 · 被引用 19 次
- Making Non-Stochastic Control (Almost) as Easy as StochasticMax SimchowitzNeurIPS 2020 · 被引用 44 次
- Leveraging Predictions in Smoothed Online Convex Optimization via Gradient-based AlgorithmsYingying Li, Na LiNeurIPS 2020 · 被引用 30 次
- Logarithmic Regret for Adversarial Online ControlDylan J. Foster, Max SimchowitzICML 2020 · 被引用 82 次
- Online learning with dynamics: A minimax perspectiveKush Bhatia, Karthik SridharanNeurIPS 2020 · 被引用 18 次
