Learning Human-like Representations to Enable Learning Human Values
Andrea Wynn, Ilia Sucholutsky, Tom Griffiths
摘要
How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of representational alignment between humans and AI agents on learning human values. Making AI systems learn human-like representations of the world has many known benefits, including improving generalization, robustness to domain shifts, and few-shot learning performance. We demonstrate that this kind of representational alignment can also support safely learning and exploring human values in the context of personalization. We begin with a theoretical prediction, show that it applies to learning human morality judgments, then show that our results generalize to ten different aspects of human values -- including ethics, honesty, and fairness -- training AI agents on each set of values in a multi-armed bandit setting, where rewards reflect human value judgments over the chosen action. Using a set of textual action descriptions, we collect value judgments from humans, as well as similarity judgments from both humans and multiple language models, and demonstrate that representational alignment enables both safe exploration and improved generalization when learning human values.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan 等ICLR 2024 · 被引用 212 次
- Improving neural network representations using human similarity judgmentsLukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen 等NeurIPS 2023 · 被引用 61 次
- Making Monolingual Sentence Embeddings Multilingual using Knowledge DistillationNils Reimers, Iryna GurevychEMNLP 2020 · 被引用 54 次
- Alignment with human representations supports robust few-shot learningIlia Sucholutsky, Tom GriffithsNeurIPS 2023 · 被引用 41 次
相关 Paper
- MAP: Multi-Human-Value Alignment PaletteXinran Wang, Qi Le, Ammar Ahmed, Enmao Diao 等ICLR 2025
- Training Socially Aligned Language Models on Simulated Social InteractionsRuibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang 等ICLR 2024 · 被引用 97 次
- Stay Moral and Explore: Learn to Behave Morally in Text-based GamesZijing Shi, Meng Fang, Yunqiu Xu, Ling Chen 等ICLR 2023
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert 等NeurIPS 2025 · 被引用 7 次
- Exploring the Association between Moral Foundations and Judgements of AI BehaviourJoe Brailsford, Frank Vetere, Eduardo VellosoCHI 2024 · 被引用 14 次
