Accelerated Learning with Linear Temporal Logic using Differentiable Simulation
Alper Kamil Bozkurt, Calin Belta, Ming C. Lin
Abstract
Ensuring that reinforcement learning (RL) controllers satisfy safety and reliability constraints in real-world settings remains challenging: state-avoidance and constrained Markov decision processes often fail to capture trajectory-level requirements or induce overly conservative behavior. Formal specification languages such as linear temporal logic (LTL) offer correct-by-construction objectives, yet their rewards are typically sparse, and heuristic shaping can undermine correctness. We introduce, to our knowledge, the first end-to-end framework that integrates LTL with differentiable simulators, enabling efficient gradient-based learning directly from formal specifications. Our method relaxes discrete automaton transitions via soft labeling of states, yielding differentiable rewards and state representations that mitigate the sparsity issue intrinsic to LTL while preserving objective soundness. We provide theoretical guarantees connecting Büchi acceptance to both discrete and differentiable LTL returns and derive a tunable bound on their discrepancy in deterministic and stochastic settings. Empirically, across complex, nonlinear, contact-rich continuous-control tasks, our approach substantially accelerates training and achieves up to twice the returns of discrete baselines. We further demonstrate compatibility with reward machines, thereby covering co-safe LTL and LTL without modification. By rendering automaton-based rewards differentiable, our work bridges formal methods and deep RL, enabling safe, specification-driven learning in continuous domains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63ffbe58-848f-4157-9286-c4335d72187aBuilds on18
- DiffTaichi: Differentiable Programming for Physical SimulationYuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun et al.ICLR 2020 · 479 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Learning Safe Multi-agent Control with Decentralized Neural Barrier CertificatesZengyi Qin, Kaiqing Zhang, Yuxiao Chen, Jingkai Chen et al.ICLR 2021 · 164 citations
- PlasticineLab: A Soft-Body Manipulation Benchmark with Differentiable PhysicsZhiao Huang, Yuanming Hu, Tao Du, Siyuan Zhou et al.ICLR 2021 · 164 citations
Related papers
- Automaton Constrained Q-LearningAnastasios Manganaris, Vittorio Giammarino, Ahmed H. QureshiNeurIPS 2025 · 3 citations
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
- Imitation Learning with Temporal Logic ConstraintsZining Fan, He ZhuNeurIPS 2025 · 2 citations
- Expressive Temporal Specifications for Reward MonitoringOmar Adalat, Francesco BelardinelliAAAI 2026
- Neuro-symbolic Hierarchical Learning for Long-Horizon Robotic TasksYu Huang, Ziji Wu, Zhengyi Ma, Kexin Ma et al.PLDI 2026
