Lune

AAAI2026Top-tier venue

Expressive Temporal Specifications for Reward Monitoring

Omar Adalat, Francesco Belardinelli

2026Year

Abstract

Specifying informative and dense reward functions remains a pivotal challenge in Reinforcement Learning, as it directly affects the efficiency of agent training. In this work, we harness the expressive power of quantitative Linear Temporal Logic on finite traces (LTL f [F]) to synthesize reward monitors that generate a dense stream of rewards for runtimeobservable state trajectories. By providing nuanced feedback during training, these monitors guide agents toward optimal behaviour and help mitigate the well-known issue of sparse rewards under long-horizon decision-making, which arises under the Boolean semantics dominating the current literature. Our framework is algorithm-agnostic and only relies on a state labelling function, and naturally accommodates specifying non-Markovian properties. Empirical results show that our quantitative monitors consistently subsume and, depending on the environment, outperform Boolean monitors in maximising a quantitative measure of task completion and in reducing convergence time.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cdefa869-9891-4011-9ca8-b5cb542daa5b

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines