ASP-Driven Emergency Planning for Norm Violations in Reinforcement Learning
Sebastian P. Adam, Thomas Eiter
Abstract
Reinforcement learning is a widely used approach for training an agent to maximize rewards in a given environment. Action policies learned with this technique see a broad range of applications in practical areas like games, healthcare, robotics, or autonomous driving. However, enforcing ethical behavior or norms based on deontic constraints that the agent should adhere to during policy execution remains a complex challenge. Especially constraints that emerge after the training can necessitate to redo policy learning, which can be costly and, more critically, time-intense. In order to mitigate this problem, we present a framework for policy fixing in case of a norm violation, which allows the agent to stay operational. Based on answer set programming (ASP), emergency plans are generated that exclude or minimize cost of norm violations by future actions in a horizon of interest. By combining and developing optimization techniques, efficient policy fixing under real-time constraints can be achieved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 174b0152-63cf-40b3-8d7e-42d7fd55ba54Related papers
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 238 citations
- Iterative Reachability Estimation for Safe Reinforcement LearningMilan Ganai, Zheng Gong, Chenning Yu, Sylvia L. Herbert et al.NeurIPS 2023 · 56 citations
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause et al.NeurIPS 2020 · 109 citations
- Mitigating Adversarial Norm Training with Moral AxiomsTaylor Olson, Kenneth D. ForbusAAAI 2023 · 7 citations
