Automating the Refinement of Reinforcement Learning Specifications
Tanmay Ambadkar, Djordje Zikelic, Abhinav Verma
Abstract
Logical specifications have been shown to help reinforcement learning algorithms in achieving complex tasks. However, when a task is under-specified, agents might fail to learn useful policies. In this work, we explore the possibility of improving coarse-grained logical specifications via an exploration-guided strategy. We propose AutoSpec, a framework that searches for a logical specification refinement whose satisfaction implies satisfaction of the original specification, but which provides additional guidance therefore making it easier for reinforcement learning algorithms to learn useful policies. AutoSpec is applicable to reinforcement learning tasks specified via the SpectRL specification logic. We exploit the compositional nature of specifications written in SpectRL, and design four refinement procedures that modify the abstract graph of the specification by either refining its existing edge specifications or by introducing new edge specifications. We prove that all four procedures maintain specification soundness, i.e. any trajectory satisfying the refined specification also satisfies the original. We then show how AutoSpec can be integrated with existing reinforcement learning algorithms for learning policies from logical specifications. Our experiments demonstrate that AutoSpec yields promising improvements in terms of the complexity of control tasks that can be solved, when refined logical specifications produced by AutoSpec are utilized. https://ambadkar.com/autospec
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee3e95c9-73f5-481f-874c-726bf7365fbeBuilds on10
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 112 citations
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho et al.NeurIPS 2021 · 107 citations
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 106 citations
- Learning Control Policies for Stochastic Systems with Reach-Avoid GuaranteesDorde Zikelic, Mathias Lechner, Thomas A. Henzinger, Krishnendu ChatterjeeAAAI 2023 · 50 citations
- Settling the Reward HypothesisMichael Bowling, John D. Martin, David Abel, Will DabneyICML 2023 · 47 citations
Related papers
- Compositional Policy Learning in Stochastic Control Systems with Formal GuaranteesDorde Zikelic, Mathias Lechner, Abhinav Verma, Krishnendu Chatterjee et al.NeurIPS 2023 · 31 citations
- HypRL: Reinforcement Learning of Control Policies for HyperpropertiesTzu-Han Hsu, Arshia Rafieioskouei, Borzoo BonakdarpourNeurIPS 2025 · 5 citations
- Automaton Constrained Q-LearningAnastasios Manganaris, Vittorio Giammarino, Ahmed H. QureshiNeurIPS 2025 · 3 citations
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda et al.AAAI 2020 · 37 citations
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 44 citations
