Reinforcement Learning with Non-Markovian Rewards
Maor Gaon, Ronen I. Brafman
Abstract
The standard RL world model is that of a Markov Decision Process (MDP). A basic premise of MDPs is that the rewards depend on the last state and action only. Yet, many real-world rewards are non-Markovian. For example, a reward for bringing coffee only if requested earlier and not yet served, is non-Markovian if the state only records current requests and deliveries. Past work considered the problem of modeling and solving MDPs with non-Markovian rewards (NMR), but we know of no principled approaches for RL with NMR. Here, we address the problem of policy learning from experience with such rewards. We describe and evaluate empirically four combinations of the classical RL algorithm Q-learning and R-max with automata learning algorithms to obtain new RL algorithms for domains with NMR. We also prove that some of these variants converge to an optimal policy in the limit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 077e0366-2508-46c2-9ee5-59ff369517dbCited by top-tier papers17
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham et al.AAAI 2021 · 62 citations
- Universal Trading for Order Execution with Oracle Policy DistillationYuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou et al.AAAI 2021 · 52 citations
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu et al.AAAI 2021 · 40 citations
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 37 citations
- Dynamic Automaton-Guided Reward Shaping for Monte Carlo Tree SearchAlvaro Velasquez, Brett Bissey, Lior Barak, Andre Beckus et al.AAAI 2021 · 23 citations
Related papers
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 4 citations
- Provably Efficient Offline Reinforcement Learning in Regular Decision ProcessesRoberto Cipollone, Anders Jonsson, Alessandro Ronca, Mohammad Sadegh TalebiNeurIPS 2023 · 7 citations
- Offline RL in Regular Decision Processes: Sample Efficiency via Language MetricsAhana Deb, Roberto Cipollone, Anders Jonsson, Alessandro Ronca et al.ICLR 2025
- Online Learning in CMDPs: Handling Stochastic and Adversarial ConstraintsFrancesco Emanuele Stradi, Jacopo Germano, Gianmarco Genalti, Matteo Castiglioni et al.ICML 2024 · 7 citations
- Reinforcement Learning When All Actions Are Not Always AvailableYash Chandak, Georgios Theocharous, Blossom Metevier, Philip S. ThomasAAAI 2020 · 8 citations
