Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement Learning
Pierriccardo Olivieri, Fausto Lasca, Alessandro Gianola, Matteo Papini
摘要
In this work, we propose a novel framework for the logical specification of non-Markovian rewards in Markov Decision Processes (MDPs) with large state spaces. Our approach leverages Linear Temporal Logic Modulo Theories over finite traces (LTLfMT), a more expressive extension of classical temporal logic in which predicates are first-order formulas of arbitrary first-order theories rather than simple Boolean variables. This enhanced expressiveness enables the specification of complex tasks over unstructured and heterogeneous data domains, promoting a unified and reusable framework that eliminates the need for manual predicate encoding. However, the increased expressive power of LTLfMT introduces additional theoretical and computational challenges compared to standard LTLf specifications. We address these challenges from a theoretical standpoint, identifying a fragment of LTLfMT that is tractable but sufficiently expressive for reward specification in an infinite-state-space context. From a practical perspective, we introduce a method based on reward machines and Hindsight Experience Replay (HER) to translate first-order logic specifications and address reward sparsity. We evaluate this approach to a continuous-control setting using Non-Linear Arithmetic Theory, showing that it enables natural specification of complex tasks. Experimental results show how a tailored implementation of HER is fundamental in solving tasks with complex goals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksYuqian Jiang, Suda Bharadwaj, Bo Wu, Rishi Shah 等AAAI 2021 · 被引用 54 次
- The Logical Options FrameworkBrandon Araki, Xiao Li, Kiran Vodrahalli, Jonathan A. DeCastro 等ICML 2021 · 被引用 44 次
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor 等NeurIPS 2024 · 被引用 19 次
- Adaptive Reactive Synthesis for LTL and LTLf Modulo TheoriesAndoni Rodríguez, César SánchezAAAI 2024 · 被引用 19 次
- Boolean Abstractions for Realizability Modulo TheoriesAndoni Rodríguez, César SánchezCAV 2023 · 被引用 18 次
相关 Paper
- Eventual Discounting Temporal Logic Counterfactual Experience ReplayCameron Voloshin, Abhinav Verma, Yisong YueICML 2023 · 被引用 24 次
- Expressive Temporal Specifications for Reward MonitoringOmar Adalat, Francesco BelardinelliAAAI 2026
- First-Order Representation Languages for Goal-Conditioned RLSimon Ståhlberg, Hector GeffnerAAAI 2026 · 被引用 1 次
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm 等ICLR 2024 · 被引用 3 次
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
