Mitigating Adversarial Norm Training with Moral Axioms
Taylor Olson, Kenneth D. Forbus
摘要
This paper addresses the issue of adversarial attacks on ethical AI systems. We investigate using moral axioms and rules of deontic logic in a norm learning framework to mitigate adversarial norm training. This model of moral intuition and construction provides AI systems with moral guard rails yet still allows for learning conventions. We evaluate our approach by drawing inspiration from a study commonly used in moral development research. This questionnaire aims to test an agent's ability to reason to moral conclusions despite opposed testimony. Our findings suggest that our model can still correctly evaluate moral situations and learn conventions in an adversarial training environment. We conclude that adding axiomatic moral prohibitions and deontic inference rules to a norm learning model makes it less vulnerable to adversarial attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than OutcomesYu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko 等ICLR 2026 · 被引用 23 次
- NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural LogicZi'ou Zheng, Xiaodan ZhuACL 2023 · 被引用 4 次
- The Moral Integrity Corpus: A Benchmark for Ethical Dialogue SystemsCaleb Ziems, Jane A. Yu, Yi-Chia Wang, Alon Y. Halevy 等ACL 2022 · 被引用 127 次
- Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language ModelsChenchen Yuan, Zheyu Zhang, Gjergji KasneciACL 2026
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentZhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal 等NeurIPS 2022 · 被引用 146 次
