Distillation of RL Policies with Formal Guarantees via Variational Abstraction of Markov Decision Processes
Florent Delgrange, Ann Nowé, Guillermo A. Pérez
Abstract
We consider the challenge of policy simplification and verification in the context of policies learned through reinforcement learning (RL) in continuous environments. In wellbehaved settings, RL algorithms have convergence guarantees in the limit. While these guarantees are valuable, they are insufficient for safety-critical applications. Furthermore, they are lost when applying advanced techniques such as deep-RL. To recover guarantees when applying advanced RL algorithms to more complex environments with (i) reachability, (ii) safety-constrained reachability, or (iii) discounted-reward objectives, we build upon the DeepMDP framework introduced by Gelada et al. to derive new bisimulation bounds between the unknown environment and a learned discrete latent model of it. Our bisimulation bounds enable the application of formal methods for Markov decision processes. Finally, we show how one can use a policy obtained via state-of-the-art RL to efficiently train a variational autoencoder that yields a discrete latent model with provably approximately correct bisimulation guarantees. Additionally, we obtain a distilled version of the policy for the latent model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2514dd1d-96e0-46b4-b647-d3c09fc2bcaeCited by top-tier papers5
- Interventionally Consistent Surrogates for Complex Simulation ModelsJoel Dyer, Nicholas Bishop, Yorgos Felekis, Fabio Massimo Zennaro et al.NeurIPS 2024 · 12 citations
- The Wasserstein Believer: Learning Belief Updates for Partially Observable Environments through Reliable Latent Space ModelsRaphaël Avalos, Florent Delgrange, Ann Nowé, Guillermo A. Pérez et al.ICLR 2024 · 10 citations
- Model-Based Offline Reinforcement Learning with Local MisspecificationKefan Dong, Yannis Flet-Berliac, Allen Nie, Emma BrunskillAAAI 2023 · 6 citations
- Deep SPI: Safe Policy Improvement via World ModelsFlorent Delgrange, Raphaël Avalos, Willem RöpkeICLR 2026 · 4 citations
- Wasserstein Auto-encoded MDPs: Formal Verification of Efficiently Distilled RL Policies with Many-sided GuaranteesFlorent Delgrange, Ann Nowé, Guillermo A. PérezICLR 2023
Builds on4
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Collapsed Amortized Variational Inference for Switching Nonlinear Dynamical SystemsZhe Dong, Bryan A. Seybold, Kevin Murphy, Hung H. BuiICML 2020 · 37 citations
- Steady State Analysis of Episodic Reinforcement LearningBojun HuangNeurIPS 2020 · 29 citations
- Global PAC Bounds for Learning Discrete Time Markov ChainsHugo Bazille, Blaise Genest, Cyrille Jégourel, Jun SunCAV 2020 · 11 citations
Related papers
- SimSR: Simple Distance-Based State Representations for Deep Reinforcement LearningHongyu Zang, Xin Li, Mingzhong WangAAAI 2022 · 20 citations
- Towards Robust Bisimulation Metric LearningMete Kemertas, Tristan Aumentado-ArmstrongNeurIPS 2021 · 68 citations
- Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real DeploymentGengyue Han, Yiheng FengICML 2026
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
- Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement LearningPrajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody H. FlemingICLR 2025
