Reinforcement Learning with Stochastic Reward Machines
Jan Corazza, Ivan Gavran, Daniel Neider
Abstract
Reward machines are an established tool for dealing with reinforcement learning problems in which rewards are sparse and depend on complex sequences of actions. However, existing algorithms for learning reward machines assume an overly idealized setting where rewards have to be free of noise. To overcome this practical limitation, we introduce a novel type of reward machines, called stochastic reward machines, and an algorithm for learning them. Our algorithm, based on constraint solving, learns minimal stochastic reward machines from the explorations of a reinforcement learning agent. This algorithm can easily be paired with existing reinforcement learning algorithms for reward machines and guarantees to converge to an optimal policy in the limit. We demonstrate the effectiveness of our algorithm in two case studies and show that it outperforms both existing methods and a naive approach for handling noisy reward functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2f708fe-a59a-4762-9a17-a72d2ab8af02Cited by top-tier papers6
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor et al.NeurIPS 2024 · 19 citations
- Hierarchies of Reward MachinesDaniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda et al.ICML 2023 · 15 citations
- Automata Learning from Preference and Equivalence QueriesEric Hsiung, Joydeep Biswas, Swarat ChaudhuriCAV 2025 · 1 citation
- Computably Continuous Reinforcement-Learning Objectives Are PAC-LearnableCambridge Yang, Michael Littman, Michael CarbinAAAI 2023
- The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed RewardsXi Chen, Zhihui Zhu, Andrew PerraultAAAI 2025
Builds on6
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 161 citations
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 96 citations
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham et al.AAAI 2021 · 62 citations
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu et al.AAAI 2021 · 40 citations
- Induction of Subgoal Automata for Reinforcement LearningDaniel Furelos-Blanco, Mark Law, Alessandra Russo, Krysia Broda et al.AAAI 2020 · 37 citations
Related papers
- Learning to Shape Rewards Using a Game of Two PartnersDavid Mguni, Taher Jafferjee, Jianhong Wang, Nicolas Perez Nieves et al.AAAI 2023 · 17 citations
- Beyond Optimism: Exploration With Partially Observable RewardsSimone Parisi, Alireza Kazemipour, Michael BowlingNeurIPS 2024 · 7 citations
- Generative Exploration and ExploitationJiechuan Jiang, Zongqing LuAAAI 2020 · 6 citations
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 2 citations
- Efficient Reinforcement Learning in Probabilistic Reward MachinesXiaofeng Lin, Xuezhou ZhangAAAI 2025 · 3 citations
