Hierarchies of Reward Machines
Daniel Furelos-Blanco, Mark Law, Anders Jonsson, Krysia Broda, Alessandra Russo
Abstract
Reward machines (RMs) are a recent formalism for representing the reward function of a reinforcement learning task through a finite-state machine whose edges encode subgoals of the task using high-level events. The structure of RMs enables the decomposition of a task into simpler and independently solvable subtasks that help tackle long-horizon and/or sparse reward tasks. We propose a formalism for further abstracting the subtask structure by endowing an RM with the ability to call other RMs, thus composing a hierarchy of RMs (HRM). We exploit HRMs by treating each call to an RM as an independently solvable subtask using the options framework, and describe a curriculum-based method to learn HRMs from traces observed by the agent. Our experiments reveal that exploiting a handcrafted HRM leads to faster convergence than with a flat HRM, and that learning an HRM is feasible in cases where its equivalent flat representation is not.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7aa586e6-5879-4175-a91e-d495267ae2c1Cited by top-tier papers3
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor et al.NeurIPS 2024 · 19 citations
- Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided SearchMax Liu, Chan-Hung Yu, Wei-Hsu Lee, Cheng-Wei Hung et al.ICLR 2025
- Beyond Fixed Tasks: Unsupervised Environment Design for Task-Level PairsDaniel Furelos-Blanco, Charles Pert, Frederik Kelbel, Alex F. Spies et al.AAAI 2026
Builds on6
- Prioritized Level ReplayMinqi Jiang, Edward Grefenstette, Tim RocktäschelICML 2021 · 211 citations
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 96 citations
- A Boolean Task Algebra for Reinforcement LearningGeraud Nangue Tasse, Steven James, Benjamin RosmanNeurIPS 2020 · 71 citations
- The Logical Options FrameworkBrandon Araki, Xiao Li, Kiran Vodrahalli, Jonathan A. DeCastro et al.ICML 2021 · 44 citations
- Reinforcement Learning with Stochastic Reward MachinesJan Corazza, Ivan Gavran, Daniel NeiderAAAI 2022 · 37 citations
Related papers
- ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement LearningRoger Creus Castanyer, Faisal Mohamed, Pablo Samuel Castro, Cyrus Neary et al.ICLR 2026 · 6 citations
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 2 citations
- Planning with Abstract Learned Models While Learning Transferable SubtasksJohn Winder, Stephanie Milani, Matthew Landen, Erebus Oh et al.AAAI 2020 · 10 citations
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari et al.NeurIPS 2025 · 5 citations
