Mitigating Adversarial Norm Training with Moral Axioms
Taylor Olson, Kenneth D. Forbus
Abstract
This paper addresses the issue of adversarial attacks on ethical AI systems. We investigate using moral axioms and rules of deontic logic in a norm learning framework to mitigate adversarial norm training. This model of moral intuition and construction provides AI systems with moral guard rails yet still allows for learning conventions. We evaluate our approach by drawing inspiration from a study commonly used in moral development research. This questionnaire aims to test an agent's ability to reason to moral conclusions despite opposed testimony. Our findings suggest that our model can still correctly evaluate moral situations and learn conventions in an adversarial training environment. We conclude that adding axiomatic moral prohibitions and deontic inference rules to a norm learning model makes it less vulnerable to adversarial attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than OutcomesYu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko et al.ICLR 2026 · 23 citations
- NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural LogicZi'ou Zheng, Xiaodan ZhuACL 2023 · 4 citations
- The Moral Integrity Corpus: A Benchmark for Ethical Dialogue SystemsCaleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy et al.ACL 2022 · 127 citations
- Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language ModelsChenchen Yuan, Zheyu Zhang, Gjergji KasneciACL 2026
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentZhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal et al.NeurIPS 2022 · 146 citations
