How to talk so AI will learn: Instructions, descriptions, and autonomy
Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths, Dylan Hadfield-Menell
Abstract
From the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a central goal of value alignment. Yet today, we lack computational models explaining such language use. To address this challenge, we formalize learning from language in a contextual bandit setting and ask how a human might communicate preferences over behaviors. We study two distinct types of language: , which provide information about the desired policy, and , which provide information about the reward function. We show that the agent's degree of autonomy determines which form of language is optimal: instructions are better in low-autonomy settings, but descriptions are better when the agent will need to act independently. We then define a pragmatic listener agent that robustly infers the speaker's reward function by reasoning about the speaker expresses themselves. We validate our models with a behavioral experiment, demonstrating that (1) our speaker model predicts human behavior, and (2) our pragmatic listener successfully recovers humans' reward functions. Finally, we show that this form of social learning can integrate with and reduce regret in traditional reinforcement learning. We hope these insights facilitate a shift from developing agents that language to agents that from it.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 139806c3-cbc3-4d37-9d75-0a0743bbcffeCited by top-tier papers4
- What Makes a Good Natural Language Prompt?Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi et al.ACL 2025 · 13 citations
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 7 citations
- M³HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed QualityZiyan Wang, Zhicheng Zhang, Fei Fang, Yali DuICML 2025
- Reward Learning from Multiple Feedback TypesYannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-AssadyICLR 2025
Builds on13
- Reward-rational (implicit) choice: A unifying formalism for reward learningHong Jun Jeon, Smitha Milli, Anca D. DraganNeurIPS 2020 · 219 citations
- Language as a Cognitive Tool to Imagine Goals in Curiosity Driven ExplorationCédric Colas, Tristan Karch, Nicolas Lair, Jean-Michel Dussoux et al.NeurIPS 2020 · 139 citations
- Skill Induction and Planning with Latent LanguagePratyusha Sharma, Antonio Torralba, Jacob AndreasACL 2022 · 127 citations
- On the Expressivity of Markov RewardDavid Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho et al.NeurIPS 2021 · 107 citations
- Semantic Exploration from Language Abstractions and Pretrained RepresentationsAllison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy et al.NeurIPS 2022 · 85 citations
Related papers
- Inferring Rewards from Language in ContextJessy Lin, Daniel Fried, Dan Klein, Anca D. DraganACL 2022 · 71 citations
- Contrastive Preference Learning: Learning from Human Feedback without Reinforcement LearningJoey Hejna, Rafael Rafailov, Harshit Sikchi, Chelsea Finn et al.ICLR 2024 · 37 citations
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan et al.AAAI 2021 · 67 citations
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 21 citations
- A Regret Minimization Framework on Preference Learning in Large Language ModelsSuhwan Kim, Taehyun Cho, Youngsoo Jang, Geon-Hyeong Kim et al.ICML 2026
