Expressive Temporal Specifications for Reward Monitoring
Omar Adalat, Francesco Belardinelli
Abstract
Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we harness the expressive power of quantitative Linear Temporal Logic on finite traces (LTL f [F]) to synthesize reward monitors that generate a dense stream of rewards for runtimeobservable state trajectories. By providing nuanced feedback during training, these monitors guide agents toward optimal behaviour and help mitigate the well-known issue of sparse rewards under long-horizon decision-making, which arises under the Boolean semantics dominating the current literature. Our framework is algorithm-agnostic and only relies on a state labelling function, and naturally accommodates specifying non-Markovian properties. Empirical results show that our quantitative monitors consistently subsume and, depending on the environment, outperform Boolean monitors in maximising a quantitative measure of task completion and in reducing convergence time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdefa869-9891-4011-9ca8-b5cb542daa5bBuilds on2
- Interpretable Reward Redistribution in Reinforcement Learning: A Causal ApproachYudi Zhang, Yali Du, Biwei Huang, Ziyan Wang et al.NeurIPS 2023 · 32 citations
- Policy Synthesis and Reinforcement Learning for Discounted LTLRajeev Alur, Osbert Bastani, Kishor Jothimurugan, Mateo Perez et al.CAV 2023 · 3 citations
Related papers
- Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement LearningPierriccardo Olivieri, Fausto Lasca, Alessandro Gianola, Matteo PapiniAAAI 2026 · 2 citations
- Accelerated Learning with Linear Temporal Logic using Differentiable SimulationAlper Kamil Bozkurt, Calin Belta, Ming C. LinICLR 2026 · 2 citations
- Learning to Follow Instructions in Text-Based GamesMathieu Tuli, Andrew C. Li, Pashootan Vaezipoor, Toryn Q. Klassen et al.NeurIPS 2022 · 21 citations
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm et al.ICLR 2024 · 3 citations
- Control Synthesis of Cyber-Physical Systems for Real-Time Specifications Through Causation-Guided Reinforcement LearningXiaochen Tang, Zhenya Zhang, Miaomiao Zhang, Jie AnRTSS 2025 · 1 citation
