Reinforcement Learning with Non-Markovian Rewards
Maor Gaon, Ronen I. Brafman
摘要
The standard RL world model is that of a Markov Decision Process (MDP). A basic premise of MDPs is that the rewards depend on the last state and action only. Yet, many real-world rewards are non-Markovian. For example, a reward for bringing coffee only if requested earlier and not yet served, is non-Markovian if the state only records current requests and deliveries. Past work considered the problem of modeling and solving MDPs with non-Markovian rewards (NMR), but we know of no principled approaches for RL with NMR. Here, we address the problem of policy learning from experience with such rewards. We describe and evaluate empirically four combinations of the classical RL algorithm Q-learning and R-max with automata learning algorithms to obtain new RL algorithms for domains with NMR. We also prove that some of these variants converge to an optimal policy in the limit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham 等AAAI 2021 · 被引用 62 次
- Universal Trading for Order Execution with Oracle Policy DistillationYuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou 等AAAI 2021 · 被引用 52 次
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu 等AAAI 2021 · 被引用 40 次
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 被引用 37 次
- Dynamic Automaton-Guided Reward Shaping for Monte Carlo Tree SearchAlvaro Velasquez, Brett Bissey, Lior Barak, Andre Beckus 等AAAI 2021 · 被引用 23 次
相关 Paper
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 被引用 4 次
- Provably Efficient Offline Reinforcement Learning in Regular Decision ProcessesRoberto Cipollone, Anders Jonsson, Alessandro Ronca, Mohammad Sadegh TalebiNeurIPS 2023 · 被引用 7 次
- Offline RL in Regular Decision Processes: Sample Efficiency via Language MetricsAhana Deb, Roberto Cipollone, Anders Jonsson, Alessandro Ronca 等ICLR 2025
- Online Learning in CMDPs: Handling Stochastic and Adversarial ConstraintsFrancesco Emanuele Stradi, Jacopo Germano, Gianmarco Genalti, Matteo Castiglioni 等ICML 2024 · 被引用 7 次
- Reinforcement Learning When All Actions Are Not Always AvailableYash Chandak, Georgios Theocharous, Blossom Metevier, Philip S. ThomasAAAI 2020 · 被引用 8 次
