Learning Safe Control via On-the-Fly Bandit Exploration
Alexandre Capone, Ryan Kazuo Cosner, Aaron D. Ames, Sandra Hirche
摘要
Control tasks with safety requirements under high levels of model uncertainty are increasingly common. Machine learning techniques are frequently used to address such tasks, typically by leveraging model error bounds to specify robust constraintbased safety filters. However, if the learned model uncertainty is very high, the corresponding filters are potentially invalid, meaning no control input satisfies the constraints imposed by the safety filter. While most works address this issue by assuming some form of safe backup controller, ours tackles it by collecting additional data on the fly using a Gaussian process bandit-type algorithm. We combine a control barrier function with a learned model to specify a robust certificate that ensures safety if feasible. Whenever infeasibility occurs, we leverage the control barrier function to guide exploration, ensuring the collected data contributes toward the closed-loop system safety. By combining a safety filter with exploration in this manner, our method provably achieves safety in a setting that allows for a zero-mean prior dynamics model, without requiring a backup controller. To the best of our knowledge, it is the first safe learning-based control method that achieves this.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Practical and Rigorous Uncertainty Bounds for Gaussian Process RegressionChristian Fiedler, Carsten W. Scherer, Sebastian TrimpeAAAI 2021 · 被引用 92 次
- Gaussian Process Uniform Error Bounds with Unknown Hyperparameters for Safety-Critical ApplicationsAlexandre Capone, Armin Lederer, Sandra HircheICML 2022 · 被引用 26 次
- Near-Optimal Multi-Agent Learning for Safe Coverage ControlManish Prajapat, Matteo Turchetta, Melanie N. Zeilinger, Andreas KrauseNeurIPS 2022 · 被引用 23 次
相关 Paper
- Safely Learning Controlled Stochastic DynamicsLuc Brogat-Motte, Alessandro Rudi, Riccardo BonalliNeurIPS 2025 · 被引用 2 次
- ActSafe: Active Exploration with Safety Constraints for Reinforcement LearningYarden As, Bhavya Sukhija, Lenart Treven, Carmelo Sferrazza 等ICLR 2025
- Neural Control and Certificate Repair via Runtime MonitoringEmily Yu, Dorde Zikelic, Thomas A. HenzingerAAAI 2025 · 被引用 5 次
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang 等ICML 2023 · 被引用 81 次
- Safety Guarantees for Neural Network Dynamic Systems via Stochastic Barrier FunctionsRayan Mazouz, Karan Muvvala, Akash Ratheesh, Luca Laurenti 等NeurIPS 2022 · 被引用 44 次
