Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W. Lee, Richard Ren, Long Phan, Norman Mu, Oliver Zhang, Dan Hendrycks
Abstract
As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tracking the emergence of goals and values has proven a longstanding problem, and despite much interest over the years it remains unclear whether current AIs have meaningful values. We propose a solution to this problem, leveraging the framework of utility functions to study the internal coherence of AI preferences. Surprisingly, we find that independently-sampled preferences in current LLMs exhibit high degrees of structural coherence, and moreover that this emerges with scale. These findings suggest that value systems emerge in LLMs in a meaningful sense, a finding with broad implications. To study these emergent value systems, we propose utility engineering as a research agenda, comprising both the analysis and control of AI utilities. We uncover problematic and often shocking values in LLM assistants despite existing control measures. These include cases where AIs value themselves over humans and are anti-aligned with specific individuals. To constrain these emergent value systems, we propose methods of utility control. As a case study, we show how aligning utilities with a citizen assembly reduces political biases and generalizes to new scenarios. Whether we like it or not, value systems have already emerged in AIs, and much work remains to fully understand and control these emergent representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91cb9b6d-9f24-4fac-84f6-98f479ac73caCited by top-tier papers12
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 296 citations
- Generative Value Conflicts Reveal LLM PrioritiesAndy Liu, Kshitish Ghate, Mona T. Diab, Daniel Fried et al.ICLR 2026 · 17 citations
- LitmusValues: Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmasYu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi et al.ICLR 2026 · 13 citations
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert et al.NeurIPS 2025 · 7 citations
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 5 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- Aligning Large Language Models with Human Preferences through Representation EngineeringWenhao Liu, Xiaohua Wang, Muling Wu, Tianlong Li et al.ACL 2024 · 4 citations
- Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypesSullam Jeoung, Yubin Ge, Haohan Wang, Jana DiesnerICLR 2025
- Do LLMs have Consistent Values?Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson et al.ICLR 2025
- Implicit Values Embedded in How Humans and LLMs Complete Subjective Everyday TasksArjun Arunasalam, Madison Pickering, Z. Berkay Celik, Blase UrEMNLP 2025
- Controllable Preference Optimization: Toward Controllable Multi-Objective AlignmentYiju Guo, Ganqu Cui, Lifan Yuan, Ning Ding et al.EMNLP 2024 · 9 citations
