DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement Learning
Mohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham, Daniel Kroening
摘要
This paper proposes DeepSynth, a method for effective training of deep Reinforcement Learning (RL) agents when the reward is sparse and non-Markovian, but at the same time progress towards the reward requires achieving an unknown sequence of high-level objectives. Our method employs a novel algorithm for synthesis of compact automata to uncover this sequential structure automatically. We synthesise a human-interpretable automaton from trace data collected by exploring the environment. The state space of the environment is then enriched with the synthesised automaton so that the generation of a control policy by deep RL is guided by the discovered structure encoded in the automaton. The proposed approach is able to cope with both high-dimensional, low-level features and unknown sparse non-Markovian rewards. We have evaluated DeepSynth's performance in a set of experiments that includes the Atari game Montezuma's Revenge. Compared to existing approaches, we obtain a reduction of two orders of magnitude in the number of iterations required for policy synthesis, and also a significant improvement in scalability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 被引用 37 次
- GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisYushi Cao, Zhiming Li, Tianpei Yang, Hao Zhang 等NeurIPS 2022 · 被引用 23 次
- Provably Efficient Offline Reinforcement Learning in Regular Decision ProcessesRoberto Cipollone, Anders Jonsson, Alessandro Ronca, Mohammad Sadegh TalebiNeurIPS 2023 · 被引用 7 次
- SYMBXRL: Symbolic Explainable Deep Reinforcement Learning for Mobile NetworksAbhishek Duttagupta, MohammadErfan Jabbari, Claudio Fiandrino, Marco Fiore 等INFOCOM 2025 · 被引用 6 次
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper5
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
- SegSort: Segmentation by Discriminative Sorting of SegmentsJyh-Jing Hwang, Stella X. Yu, Jianbo Shi, Maxwell D. Collins 等ICCV 2019 · 被引用 160 次
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 被引用 96 次
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda 等AAAI 2020 · 被引用 37 次
- Learning Concise Models from Long Execution TracesNatasha Yogananda Jeppu, Thomas F. Melham, Daniel Kroening, John O'LearyDAC 2020
相关 Paper
- Dynamic Automaton-Guided Reward Shaping for Monte Carlo Tree SearchAlvaro Velasquez, Brett Bissey, Lior Barak, Andre Beckus 等AAAI 2021 · 被引用 23 次
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu 等AAAI 2021 · 被引用 40 次
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor 等NeurIPS 2024 · 被引用 19 次
- PoE-World: Compositional World Modeling with Products of Programmatic ExpertsTop Piriyakulkij, Yichao Liang, Hao Tang, Adrian Weller 等NeurIPS 2025 · 被引用 31 次
- Programmatic Reinforcement Learning without OraclesWenjie Qiu, He ZhuICLR 2022 · 被引用 42 次
