Learning Human-like Representations to Enable Learning Human Values
Andrea Wynn, Ilia Sucholutsky, Tom Griffiths
Abstract
How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of representational alignment between humans and AI agents on learning human values. Making AI systems learn human-like representations of the world has many known benefits, including improving generalization, robustness to domain shifts, and few-shot learning performance. We demonstrate that this kind of representational alignment can also support safely learning and exploring human values in the context of personalization. We begin with a theoretical prediction, show that it applies to learning human morality judgments, then show that our results generalize to ten different aspects of human values -- including ethics, honesty, and fairness -- training AI agents on each set of values in a multi-armed bandit setting, where rewards reflect human value judgments over the chosen action. Using a set of textual action descriptions, we collect value judgments from humans, as well as similarity judgments from both humans and multiple language models, and demonstrate that representational alignment enables both safe exploration and improved generalization when learning human values.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21b7155a-bb00-4641-998e-0929e6f331feBuilds on7
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan et al.ICLR 2024 · 212 citations
- Improving neural network representations using human similarity judgmentsLukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen et al.NeurIPS 2023 · 61 citations
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 54 citations
- Alignment with human representations supports robust few-shot learningIlia Sucholutsky, Tom GriffithsNeurIPS 2023 · 41 citations
Related papers
- MAP: Multi-Human-Value Alignment PaletteXinran Wang, Qi Le, Ammar Ahmed, Enmao Diao et al.ICLR 2025
- Training Socially Aligned Language Models on Simulated Social InteractionsRuibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang et al.ICLR 2024 · 97 citations
- Stay Moral and Explore: Learn to Behave Morally in Text-based GamesZijing Shi, Meng Fang, Yunqiu Xu, Ling Chen et al.ICLR 2023
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert et al.NeurIPS 2025 · 7 citations
- Exploring the Association between Moral Foundations and Judgements of AI BehaviourJoe Brailsford, Frank Vetere, Eduardo VellosoCHI 2024 · 14 citations
