How to talk so AI will learn: Instructions, descriptions, and autonomy
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths, Dylan Hadfield-Menell
摘要
From the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a central goal of value alignment. Yet today, we lack computational models explaining such language use. To address this challenge, we formalize learning from language in a contextual bandit setting and ask how a human might communicate preferences over behaviors. We study two distinct types of language: , which provide information about the desired policy, and , which provide information about the reward function. We show that the agent's degree of autonomy determines which form of language is optimal: instructions are better in low-autonomy settings, but descriptions are better when the agent will need to act independently. We then define a pragmatic listener agent that robustly infers the speaker's reward function by reasoning about the speaker expresses themselves. We validate our models with a behavioral experiment, demonstrating that (1) our speaker model predicts human behavior, and (2) our pragmatic listener successfully recovers humans' reward functions. Finally, we show that this form of social learning can integrate with and reduce regret in traditional reinforcement learning. We hope these insights facilitate a shift from developing agents that language to agents that from it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi 等ACL 2025 · 被引用 13 次
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 被引用 7 次
- M³HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed QualityZiyan Wang, Zhicheng Zhang, Fei Fang, Yali DuICML 2025
- Reward Learning from Multiple Feedback TypesYannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-AssadyICLR 2025
它引用的顶会 Paper13
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 被引用 219 次
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux 等NeurIPS 2020 · 被引用 139 次
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 被引用 127 次
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho 等NeurIPS 2021 · 被引用 107 次
- Semantic Exploration from Language Abstractions and Pretrained RepresentationsAllison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy 等NeurIPS 2022 · 被引用 85 次
相关 Paper
- Inferring Rewards from Language in ContextJessy Lin, Daniel Fried, Dan Klein, Anca D. DraganACL 2022 · 被引用 71 次
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn 等ICLR 2024 · 被引用 37 次
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan 等AAAI 2021 · 被引用 67 次
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim 等ICML 2026
