Reinforcement Learning with LTL and ω-Regular Objectives via Optimality-Preserving Translation to Average Rewards
Xuan-Bach Le, Dominik Wagner, Leon Witzman, Alexander Rabinovich, Luke Ong
摘要
Linear temporal logic (LTL) and, more generally, -regular objectives are alternatives to the traditional discount sum and average reward objectives in reinforcement learning (RL), offering the advantage of greater comprehensibility and hence explainability. In this work, we study the relationship between these objectives. Our main result is that each RL problem for -regular objectives can be reduced to a limit-average reward problem in an optimality-preserving fashion, via (finite-memory) reward machines. Furthermore, we demonstrate the efficacy of this approach by showing that optimal policies for limit-average problems can be found asymptotically by solving a sequence of discount-sum problems approximately. Consequently, we resolve an open problem: optimal policies for LTL and -regular objectives can be learned asymptotically.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Accelerated Learning with Linear Temporal Logic using Differentiable SimulationAlper Kamil Bozkurt, Calin Belta, Ming C. LinICLR 2026 · 被引用 2 次
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
- Reinforcement Learning for Reachability: Guaranteeing Asymptotic OptimalityAmogh Palasamudram, Jakub Svoboda, Suguman Bansal, Krishnendu ChatterjeeICML 2026
- Stochastic Minimum-Cost Reach-Avoid Reinforcement LearningJingduo Pan, Taoran Wu, Yiling Xue, Bai XueICML 2026
它引用的顶会 Paper5
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 被引用 82 次
- Policy Optimization with Linear Temporal Logic ConstraintsCameron Voloshin, Hoang Minh Le, Swarat Chaudhuri, Yisong YueNeurIPS 2022 · 被引用 28 次
- Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount FactorJulien Grand-Clément, Marek PetrikNeurIPS 2023 · 被引用 25 次
- Eventual Discounting Temporal Logic Counterfactual Experience ReplayCameron Voloshin, Abhinav Verma, Yisong YueICML 2023 · 被引用 24 次
相关 Paper
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm 等ICLR 2024 · 被引用 3 次
- Policy Synthesis and Reinforcement Learning for Discounted LTLRajeev Alur, Osbert Bastani, Kishor Jothimurugan, Mateo Perez 等CAV 2023 · 被引用 3 次
- Computably Continuous Reinforcement-Learning Objectives Are PAC-LearnableCambridge Yang, Michael Littman, Michael CarbinAAAI 2023
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
- Imitation Learning with Temporal Logic ConstraintsZining Fan, He ZhuNeurIPS 2025 · 被引用 2 次
