DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement Learning
Mohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, Daniel Kroening
Abstract
This paper proposes DeepSynth, a method for effective training of deep Reinforcement Learning (RL) agents when the reward is sparse and non-Markovian, but at the same time progress towards the reward requires achieving an unknown sequence of high-level objectives. Our method employs a novel algorithm for synthesis of compact automata to uncover this sequential structure automatically. We synthesise a human-interpretable automaton from trace data collected by exploring the environment. The state space of the environment is then enriched with the synthesised automaton so that the generation of a control policy by deep RL is guided by the discovered structure encoded in the automaton. The proposed approach is able to cope with both high-dimensional, low-level features and unknown sparse non-Markovian rewards. We have evaluated DeepSynth's performance in a set of experiments that includes the Atari game Montezuma's Revenge. Compared to existing approaches, we obtain a reduction of two orders of magnitude in the number of iterations required for policy synthesis, and also a significant improvement in scalability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ada2fb08-e8bf-441d-a440-afd814b18d99Cited by top-tier papers8
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 37 citations
- GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisYushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang et al.NeurIPS 2022 · 23 citations
- Provably Efficient Offline Reinforcement Learning in Regular Decision ProcessesRoberto Cipollone, Anders Jonsson, Alessandro Ronca, Mohammad Sadegh TalebiNeurIPS 2023 · 7 citations
- SYMBXRL: Symbolic Explainable Deep Reinforcement Learning for Mobile NetworksAbhishek Duttagupta, MohammadErfan Jabbari, Claudio Fiandrino, Marco Fiore et al.INFOCOM 2025 · 6 citations
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu et al.ICLR 2026 · 6 citations
Builds on5
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
- SegSort: Segmentation by Discriminative Sorting of SegmentsJyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins et al.ICCV 2019 · 160 citations
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 96 citations
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda et al.AAAI 2020 · 37 citations
- Learning Concise Models from Long Execution TracesNatasha Yogananda Jeppu, Thomas F. Melham, Daniel Kroening, John O'LearyDAC 2020
Related papers
- Dynamic Automaton-Guided Reward Shaping for Monte Carlo Tree SearchAlvaro Velasquez, Brett Bissey, Lior Barak, Andre Beckus et al.AAAI 2021 · 23 citations
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu et al.AAAI 2021 · 40 citations
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor et al.NeurIPS 2024 · 19 citations
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller et al.NeurIPS 2025 · 31 citations
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 42 citations
