Eventual Discounting Temporal Logic Counterfactual Experience Replay
Cameron Voloshin, Abhinav Verma, Yisong Yue
Abstract
Linear temporal logic (LTL) offers a simplified way of specifying tasks for policy optimization that may otherwise be difficult to describe with scalar reward functions. However, the standard RL framework can be too myopic to find maximally LTL satisfying policies. This paper makes two contributions. First, we develop a new value-function based proxy, using a technique we call eventual discounting, under which one can find policies that satisfy the LTL specification with highest achievable probability. Second, we develop a new experience replay method for generating off-policy data from on-policy rollouts via counterfactual reasoning on different ways of satisfying the LTL specification. Our experiments, conducted in both discrete and continuous state-action spaces, confirm the effectiveness of our counterfactual experience replay approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e12d33ad-62cf-4e64-b18c-bf019176a1a2Cited by top-tier papers8
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 44 citations
- Reinforcement Learning with LTL and ω-Regular Objectives via Optimality-Preserving Translation to Average RewardsXuan-Bach Le, Dominik Wagner, Leon Witzman, Alexander Rabinovich et al.NeurIPS 2024 · 17 citations
- One Subgoal at a Time: Zero-Shot Generalization to Arbitrary Linear Temporal Logic Requirements in Multi-Task Reinforcement LearningZijian Guo, Ilker Isik, H. M. Sabbir Ahmad, Wenchao LiNeurIPS 2025 · 13 citations
- Ground-Compose-Reinforce: Grounding Language in Agentic Behaviours using Limited DataAndrew C. Li, Toryn Q. Klassen, Andrew Wang, Parand A. Alamdari et al.NeurIPS 2025 · 5 citations
- Imitation Learning with Temporal Logic ConstraintsZining Fan, He ZhuNeurIPS 2025 · 2 citations
Builds on3
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho et al.NeurIPS 2021 · 107 citations
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 106 citations
- Policy Optimization with Linear Temporal Logic ConstraintsCameron Voloshin, Hoang Minh Le, Swarat Chaudhuri, Yisong YueNeurIPS 2022 · 28 citations
Related papers
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
- Do It for HER: First-Order Temporal Logic Reward Specification in Reinforcement LearningPierriccardo Olivieri, Fausto Lasca, Alessandro Gianola, Matteo PapiniAAAI 2026 · 2 citations
- Policy Synthesis and Reinforcement Learning for Discounted LTLRajeev Alur, Osbert Bastani, Kishor Jothimurugan, Mateo Perez et al.CAV 2023 · 3 citations
- Synthesis from Satisficing and Temporal GoalsSuguman Bansal, Lydia E. Kavraki, Moshe Y. Vardi, Andrew M. WellsAAAI 2022 · 6 citations
- Learning to Follow Instructions in Text-Based GamesMathieu Tuli, Andrew C. Li, Pashootan Vaezipoor, Toryn Q. Klassen et al.NeurIPS 2022 · 21 citations
