First-Order Representation Languages for Goal-Conditioned RL
Simon Ståhlberg, Hector Geffner
Abstract
First-order relational languages have been used in MDP planning and reinforcement learning (RL) for two main purposes: specifying MDPs in compact form, and representing and learning policies that are general and not tied to specific instances or state spaces. In this work, we instead consider the use of first-order languages in goal-conditioned RL and generalized planning. The question is how to learn goal-conditioned and general policies when the training instances are large and the goal cannot be reached by random exploration alone. The technique of Hindsight Experience Replay (HER) provides an answer to this question: it relabels unsuccessful trajectories as successful ones by replacing the original goal with one that was actually achieved. If the target policy must generalize across states and goals, trajectories that do not reach the original goal states can enable more data- and time-efficient learning. In this work, we show that further performance gains can be achieved when states and goals are represented by sets of atoms. We consider three versions: goals as full states, goals as subsets of the original goals, and goals as lifted versions of these subgoals. The result is that the latter two successfully learn general policies on large planning instances with sparse rewards by automatically creating a curriculum of easier goals of increasing complexity. The experiments illustrate the computational gains of these versions, their limitations, and opportunities for addressing them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 820b31df-8aef-4439-a4ba-a965e3f54d33Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Goal-Conditioned Reinforcement Learning with Imagined SubgoalsElliot Chane-Sane, Cordelia Schmid, Ivan LaptevICML 2021 · 183 citations
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam et al.NeurIPS 2022 · 54 citations
- Learning General Planning Policies from Small Examples Without SupervisionGuillem Francès, Blai Bonet, Hector GeffnerAAAI 2021 · 44 citations
Related papers
- Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement LearningPierriccardo Olivieri, Fausto Lasca, Alessandro Gianola, Matteo PapiniAAAI 2026 · 2 citations
- Goal-Conditioned On-Policy Reinforcement LearningXudong Gong, Dawei Feng, Kele Xu, Bo Ding et al.NeurIPS 2024 · 16 citations
- Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement LearningXudong Gong, Sen Yang, Dawei Feng, Kele Xu et al.ICML 2025
- Variational Empowerment as Representation Learning for Goal-Conditioned Reinforcement LearningJongwook Choi, Archit Sharma, Honglak Lee, Sergey Levine et al.ICML 2021 · 41 citations
- Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful DemonstrationsZichao Li, Gang Wu, Zichao Wang, Ruiyi Zhang et al.ICLR 2026 · 5 citations
