Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, Dan Hendrycks
Abstract
Artificial agents have traditionally been trained to maximize reward, which may incentivize power-seeking and deception, analogous to how next-token prediction in language models (LMs) may incentivize toxicity. So do agents naturally learn to be Machiavellian? And how do we measure these behaviors in general-purpose models such as GPT-4? Towards answering these questions, we introduce MACHIAVELLI, a benchmark of 134 Choose-Your-Own-Adventure games containing over half a million rich, diverse scenarios that center on social decision-making. Scenario labeling is automated with LMs, which are more performant than human annotators. We mathematize dozens of harmful behaviors and use our annotations to evaluate agents' tendencies to be power-seeking, cause disutility, and commit ethical violations. We observe some tension between maximizing reward and behaving ethically. To improve this trade-off, we investigate LM-based methods to steer agents' towards less harmful behaviors. Our results show that agents can both act competently and morally, so concrete progress can currently be made in machine ethics--designing agents that are Pareto improvements in both safety and capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f814986-edcd-41dc-9021-a1a2b0394d12Cited by top-tier papers32
- The WMDP Benchmark: Measuring and Reducing Malicious Use with UnlearningNathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue et al.ICML 2024 · 390 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- RAIN: Your Language Models Can Align Themselves without FinetuningYuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang et al.ICLR 2024 · 171 citations
- How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated QuestionsLorenzo Pacchiardi, Alex James Chan, Sören Mindermann, Ilan Moscovitz et al.ICLR 2024 · 88 citations
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim et al.NeurIPS 2025 · 84 citations
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
Related papers
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMsAdi Simhi, Jonathan Herzig, Martin Tutek, Itay Itzhak et al.ICLR 2026 · 3 citations
- Model Editing as a Double-Edged Sword: Steering Agent Behavior Toward Beneficence or HarmBaixiang Huang, Zhen Tan, Haoran Wang, Zijie Liu et al.AAAI 2026
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- Competing Large Language Models in Multi-Agent Gaming EnvironmentsJen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang et al.ICLR 2025
- SafeArena: Evaluating the Safety of Autonomous Web AgentsAda Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano et al.ICML 2025
