HypRL: Reinforcement Learning of Control Policies for Hyperproperties
Tzu-Han Hsu, Arshia Rafieioskouei, Borzoo Bonakdarpour
摘要
Reward shaping in multi-agent reinforcement learning (MARL) for complex tasks remains a significant challenge. Existing approaches often fail to find optimal solutions or cannot efficiently handle such tasks. We propose HYPRL, a specification-guided reinforcement learning framework that learns control policies w.r.t. hyperproperties expressed in HyperLTL. Hyperproperties constitute a powerful formalism for specifying objectives and constraints over sets of execution traces across agents. To learn policies that maximize the satisfaction of a HyperLTL formula , we apply Skolemization to manage quantifier alternations and define quantitative robustness functions to shape rewards over execution traces of a Markov decision process with unknown transitions. A suitable RL algorithm is then used to learn policies that collectively maximize the expected reward and, consequently, increase the probability of satisfying . We evaluate HYPRL on a diverse set of benchmarks, including safety-aware planning, Deep Sea Treasure, and the Post Correspondence Problem. We also compare with specification-driven baselines to demonstrate the effectiveness and efficiency of HYPRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Compositional Reinforcement Learning from Logical SpecificationsKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurNeurIPS 2021 · 被引用 112 次
- Shield Decentralization for Safe Multi-Agent Reinforcement LearningDaniel Melcer, Christopher Amato, Stavros TripakisNeurIPS 2022 · 被引用 26 次
- Don't Pour Cereal into Coffee: Differentiable Temporal Logic for Temporal Action SegmentationZiwei Xu, Yogesh S. Rawat, Yongkang Wong, Mohan S. Kankanhalli 等NeurIPS 2022 · 被引用 18 次
- Specification-Guided Learning of Nash Equilibria with High Social WelfareKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurCAV 2022 · 被引用 9 次
相关 Paper
- DeepLTL: Learning to Efficiently Satisfy Complex LTL Specifications for Multi-Task RLMathias Jackermeier, Alessandro AbateICLR 2025
- Instructing Goal-Conditioned Reinforcement Learning Agents with Temporal Logic ObjectivesWenjie Qiu, Wensen Mao, He ZhuNeurIPS 2023 · 被引用 44 次
- HMRL: Hyper-Meta Learning for Sparse Reward Reinforcement Learning ProblemYun Hua, Xiangfeng Wang, Bo Jin, Wenhao Li 等KDD 2021 · 被引用 6 次
- Automaton Constrained Q-LearningAnastasios Manganaris, Vittorio Giammarino, Ahmed H. QureshiNeurIPS 2025 · 被引用 3 次
- On the Expressivity of Objective-Specification Formalisms in Reinforcement LearningRohan Subramani, Marcus Williams, Max Heitmann, Halfdan Holm 等ICLR 2024 · 被引用 3 次
