Geometric Active Exploration in Markov Decision Processes: the Benefit of Abstraction
Riccardo De Santi, Federico Arangath Joseph, Noah Liniger, Mirco Mutti, Andreas Krause
Abstract
How can a scientist use a Reinforcement Learning (RL) algorithm to design experiments over a dynamical system's state space? In the case of finite and Markovian systems, an area called Active Exploration (AE) relaxes the optimization problem of experiments design into Convex RL, a generalization of RL admitting a wider notion of reward. Unfortunately, this framework is currently not scalable and the potential of AE is hindered by the vastness of experiment spaces typical of scientific discovery applications. However, these spaces are often endowed with natural geometries, e.g., permutation invariance in molecular design, that an agent could leverage to improve the statistical and computational efficiency of AE. To achieve this, we bridge AE and MDP homomorphisms, which offer a way to exploit known geometric structures via abstraction. Towards this goal, we make two fundamental contributions: we extend MDP homomorphisms formalism to Convex RL, and we present, to the best of our knowledge, the first analysis that formally captures the benefit of abstraction via homomorphisms on sample efficiency. Ultimately, we propose the Geometric Active Exploration (GAE) algorithm, which we analyse theoretically and experimentally in environments motivated by problems in scientific discovery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-TuningRiccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh, Zebang Shen et al.NeurIPS 2025 · 17 citations
- Verifier-Constrained Flow Expansion for Discovery Beyond the DataRiccardo De Santi, Kimon Protopapas, Ya-Ping Hsieh, Andreas KrauseICLR 2026 · 6 citations
- Efficient Tail-Aware Generative Optimization via Flow Model Fine-TuningZifan Wang, Riccardo De Santi, Xiaoyu Mo, Michael Zavlanos et al.ICML 2026 · 4 citations
- A Unified Density Operator View of Flow Control and MergingRiccardo De Santi, Malte Franke, Ya-Ping Hsieh, Andreas KrauseICML 2026 · 2 citations
- Provable Maximum Entropy Manifold Exploration via Diffusion ModelsRiccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh, Zebang Shen et al.ICML 2025
Builds on9
- MDP Homomorphic Networks: Group Symmetries in Reinforcement LearningElise van der Pol, Daniel E. Worrall, Herke van Hoof, Frans A. Oliehoek et al.NeurIPS 2020 · 203 citations
- Variational Policy Gradient Method for Reinforcement Learning with General UtilitiesJunyu Zhang, Alec Koppel, Amrit Singh Bedi, Csaba Szepesvári et al.NeurIPS 2020 · 170 citations
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 96 citations
- The Importance of Non-Markovianity in Maximum State Entropy ExplorationMirco Mutti, Riccardo De Santi, Marcello RestelliICML 2022 · 45 citations
- Active Exploration for Inverse Reinforcement LearningDavid Lindner, Andreas Krause, Giorgia RamponiNeurIPS 2022 · 36 citations
Related papers
- Kinematic State Abstraction and Provably Efficient Rich-Observation Reinforcement LearningDipendra Misra, Mikael Henaff, Akshay Krishnamurthy, John LangfordICML 2020 · 158 citations
- Boosting Sample Efficiency and Generalization in Multi-agent Reinforcement Learning via EquivarianceJoshua McClellan, Naveed Haghani, John Winder, Furong Huang et al.NeurIPS 2024 · 21 citations
- Symmetric Replay Training: Enhancing Sample Efficiency in Deep Reinforcement Learning for Combinatorial OptimizationHyeonah Kim, Minsu Kim, Sungsoo Ahn, Jinkyoo ParkICML 2024 · 9 citations
- Continuous MDP Homomorphisms and Homomorphic Policy GradientSahand Rezaei-Shoshtari, Rosie Zhao, Prakash Panangaden, David Meger et al.NeurIPS 2022 · 34 citations
- Subequivariant Graph Reinforcement Learning in 3D EnvironmentsRunfa Chen, Jiaqi Han, Fuchun Sun, Wenbing HuangICML 2023 · 14 citations
