Reward Machines for Deep RL in Noisy and Uncertain Environments
Andrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor, Rodrigo Toro Icarte, Sheila A. McIlraith
Abstract
Reward Machines provide an automaton-inspired structure for specifying instructions, safety constraints, and other temporally extended reward-worthy behaviour. By exposing the underlying structure of a reward function, they enable the decomposition of an RL task, leading to impressive gains in sample efficiency. Although Reward Machines and similar formal specifications have a rich history of application towards sequential decision-making problems, they critically rely on a ground-truth interpretation of the domain-specific vocabulary that forms the building blocks of the reward function--such ground-truth interpretations are elusive in the real world due in part to partial observability and noisy sensing. In this work, we explore the use of Reward Machines for Deep RL in noisy and uncertain environments. We characterize this problem as a POMDP and propose a suite of RL algorithms that exploit task structure under uncertain interpretation of the domain-specific vocabulary. Through theory and experiments, we expose pitfalls in naive approaches to this problem while simultaneously demonstrating how task structure can be successfully leveraged under noisy interpretations of the vocabulary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79243fd1-c3a2-433c-937e-663a5a2993a2Cited by top-tier papers5
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement LearningRoger Creus Castanyer, Faisal Mohamed, Pablo Samuel Castro, Cyrus Neary et al.ICLR 2026 · 6 citations
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques et al.ICLR 2026 · 6 citations
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari et al.NeurIPS 2025 · 5 citations
- Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement LearningPierriccardo Olivieri, Fausto Lasca, Alessandro Gianola, Matteo PapiniAAAI 2026 · 2 citations
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online VideosBowen Baker, Ilge Akkaya, Peter Zhokhov, Joost Huizinga et al.NeurIPS 2022 · 458 citations
- STEVE-1: A Generative Model for Text-to-Behavior in MinecraftShalev Lifshitz, Keiran Paster, Harris Chan, Jimmy Ba et al.NeurIPS 2023 · 123 citations
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 106 citations
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah et al.AAAI 2021 · 54 citations
Related papers
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 37 citations
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 2 citations
- RLang: A Declarative Language for Describing Partial World Knowledge to Reinforcement Learning AgentsRafael Rodríguez-Sánchez, Benjamin Adin Spiegel, Jennifer Wang, Roma Patel et al.ICML 2023 · 4 citations
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham et al.AAAI 2021 · 62 citations
- A Hierarchical Bayesian Approach to Inverse Reinforcement Learning with Symbolic Reward MachinesWeichao Zhou, Wenchao LiICML 2022 · 15 citations
