Lune

NeurIPS2022顶会

Defining and Characterizing Reward Gaming

Joar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David Krueger

出版方
2022年份
466被引次数
150顶会引用

摘要

We provide the first formal definition of reward gaming, a phenomenon where optimizing an imperfect proxy reward function, R, leads to poor performance according to a true reward function, R. We say that a proxy is ungameable if increasing the expected proxy return can never decrease the expected true return. Intuitively, it should be possible to create an ungameable proxy by overlooking fine-grained distinctions between roughly equivalent outcomes, but we show this is usually not the case. A key insight is that the linearity of reward (as a function of state-action visit counts) makes ungameability a very strong condition. In particular, for the set of all stochastic policies, two reward functions can only be ungameable if one of them is constant. We thus turn our attention to deterministic policies and finite sets of stochastic policies, where non-trivial ungameable pairs always exist, and establish necessary and sufficient conditions for the existence of simplifications, an important special case of ungameability. Our results reveal a tension between using reward functions to specify narrow tasks and aligning AI systems with human values.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4e82b84b-2869-419b-8f18-1126c53409fa

引用它的顶会 Paper150

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖