Information-Theoretic Safe Exploration with Gaussian Processes
Alessandro G. Bottero, Carlos E. Luis, Julia Vinogradska, Felix Berkenkamp, Jan Peters
Abstract
We consider a sequential decision making task where we are not allowed to evaluate parameters that violate an a priori unknown (safety) constraint. A common approach is to place a Gaussian process prior on the unknown constraint and allow evaluations only in regions that are safe with high probability. Most current methods rely on a discretization of the domain and cannot be directly extended to the continuous case. Moreover, the way in which they exploit regularity assumptions about the constraint introduces an additional critical hyperparameter. In this paper, we propose an information-theoretic safe exploration criterion that directly exploits the GP posterior to identify the most informative safe parameters to evaluate. Our approach is naturally applicable to continuous domains and does not require additional hyperparameters. We theoretically analyze the method and show that we do not violate the safety constraint with high probability and that we explore by learning about the constraint up to arbitrary precision. Empirical evaluations demonstrate improved data-efficiency and scalability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Transductive Active Learning: Theory and ApplicationsJonas Hübotter, Bhavya Sukhija, Lenart Treven, Yarden As et al.NeurIPS 2024 · 24 citations
- Beyond Noisy-TVs: Noise-Robust Exploration Via Learning Progress MonitoringZhibo Hou, Zhiyu An, Wan DuICLR 2026 · 3 citations
- Safely Learning Controlled Stochastic DynamicsLuc Brogat-Motte, Alessandro Rudi, Riccardo BonalliNeurIPS 2025 · 2 citations
- Safe Exploration in Dose Finding Clinical Trials with Heterogeneous ParticipantsIsabel Chien, Wessel P. Bruinsma, Javier González Hernández, Richard E. TurnerICML 2024
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
Builds on1
Related papers
- Gaussian Process Uniform Error Bounds with Unknown Hyperparameters for Safety-Critical ApplicationsAlexandre Capone, Armin Lederer, Sandra HircheICML 2022 · 26 citations
- Learning Safe Control via On-the-Fly Bandit ExplorationAlexandre Capone, Ryan Kazuo Cosner, Aaron D. Ames, Sandra HircheICML 2025
- Policy Search via Bayesian Optimization with Temporal Difference Gaussian ProcessesArmin Lederer, Anuj Srivastava, Marco Bagatella, Andreas KrauseICML 2026
- Safe Exploration in Reinforcement Learning: A Generalized Formulation and AlgorithmsAkifumi Wachi, Wataru Hashimoto, Xun Shen, Kazumune HashimotoNeurIPS 2023 · 38 citations
- Tree ensemble kernels for Bayesian optimization with known constraints over mixed-feature spacesAlexander Thebelt, Calvin Tsay, Robert M. Lee, Nathan Sudermann-Merx et al.NeurIPS 2022 · 18 citations
