Reinforcement Learning Under Moral Uncertainty
Adrien Ecoffet, Joel Lehman
摘要
An ambitious goal for machine learning is to create agents that behave ethically: The capacity to abide by human moral norms would greatly expand the context in which autonomous agents could be practically and safely deployed, e.g. fully autonomous vehicles will encounter charged moral decisions that complicate their deployment. While ethical agents could be trained by rewarding correct behavior under a specific moral theory (e.g. utilitarianism), there remains widespread disagreement about the nature of morality. Acknowledging such disagreement, recent work in moral philosophy proposes that ethical behavior requires acting under moral uncertainty, i.e. to take into account when acting that one's credence is split across several plausible ethical theories. This paper translates such insights to the field of reinforcement learning, proposes two training methods that realize different points among competing desiderata, and trains agents in simple environments to act under moral uncertainty. The results illustrate (1) how such uncertainty can help curb extreme behavior from commitment to single theories and (2) several technical complications arising from attempting to ground moral philosophy in RL (e.g. how can a principled trade-off between two competing but incomparable reward functions be reached). The aim is to catalyze progress towards morally-competent agents and highlight the potential of RL to contribute towards the computational grounding of moral philosophy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Flexible Attention-Based Multi-Policy Fusion for Efficient Deep Reinforcement LearningZih-Yun Chiu, Yi-Lin Tuan, William Yang Wang, Michael C. YipNeurIPS 2023 · 被引用 7 次
- Moral Uncertainty and the Problem of FanaticismJazon Szabo, Natalia Criado, Jose Such, Sanjay ModgilAAAI 2024 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Symmetric Machine Theory of MindMelanie Sclar, Graham Neubig, Yonatan BiskICML 2022 · 被引用 22 次
- Epistemic Gain, Aleatoric Cost: Uncertainty Decomposition in Multi-Agent Debate for Math ReasoningDan Qiao, Binbin Chen, Fengyu Cai, Jianlong Chen 等ICML 2026 · 被引用 3 次
- Robust Multi-Agent Reinforcement Learning with Model UncertaintyKaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc 等NeurIPS 2020 · 被引用 118 次
- Automatic Curriculum Learning through Value DisagreementYunzhi Zhang, Pieter Abbeel, Lerrel PintoNeurIPS 2020 · 被引用 132 次
- ASP-Driven Emergency Planning for Norm Violations in Reinforcement LearningSebastian P. Adam, Thomas EiterAAAI 2025
