Information-Theoretic Safe Exploration with Gaussian Processes
Alessandro G. Bottero, Carlos E. Luis, Julia Vinogradska, Felix Berkenkamp, Jan Peters
摘要
We consider a sequential decision making task where we are not allowed to evaluate parameters that violate an a priori unknown (safety) constraint. A common approach is to place a Gaussian process prior on the unknown constraint and allow evaluations only in regions that are safe with high probability. Most current methods rely on a discretization of the domain and cannot be directly extended to the continuous case. Moreover, the way in which they exploit regularity assumptions about the constraint introduces an additional critical hyperparameter. In this paper, we propose an information-theoretic safe exploration criterion that directly exploits the GP posterior to identify the most informative safe parameters to evaluate. Our approach is naturally applicable to continuous domains and does not require additional hyperparameters. We theoretically analyze the method and show that we do not violate the safety constraint with high probability and that we explore by learning about the constraint up to arbitrary precision. Empirical evaluations demonstrate improved data-efficiency and scalability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Transductive Active Learning: Theory and ApplicationsJonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As 等NeurIPS 2024 · 被引用 24 次
- Beyond Noisy-TVs: Noise-Robust Exploration Via Learning Progress MonitoringZhibo Hou, Zhiyu An, Wan DuICLR 2026 · 被引用 3 次
- Safely Learning Controlled Stochastic DynamicsLuc Brogat-Motte, Alessandro Rudi, Riccardo BonalliNeurIPS 2025 · 被引用 2 次
- Safe Exploration in Dose Finding Clinical Trials with Heterogeneous ParticipantsIsabel Chien, Wessel P. Bruinsma, Javier González Hernández, Richard E. TurnerICML 2024
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
它引用的顶会 Paper1
相关 Paper
- Gaussian Process Uniform Error Bounds with Unknown Hyperparameters for Safety-Critical ApplicationsAlexandre Capone, Armin Lederer, Sandra HircheICML 2022 · 被引用 26 次
- Learning Safe Control via On-the-Fly Bandit ExplorationAlexandre Capone, Ryan Kazuo Cosner, Aaron D. Ames, Sandra HircheICML 2025
- Policy Search via Bayesian Optimization with Temporal Difference Gaussian ProcessesArmin Lederer, Anuj Srivastava, Marco Bagatella, Andreas KrauseICML 2026
- Safe Exploration in Reinforcement Learning: A Generalized Formulation and AlgorithmsAkifumi Wachi, Wataru Hashimoto, Xun Shen, Kazumune HashimotoNeurIPS 2023 · 被引用 38 次
- Tree ensemble kernels for Bayesian optimization with known constraints over mixed-feature spacesAlexander Thebelt, Calvin Tsay, Robert M. Lee, Nathan Sudermann-Merx 等NeurIPS 2022 · 被引用 18 次
