Inferring Rewards from Language in Context
Jessy Lin, Daniel Fried, Dan Klein, Anca D. Dragan
Abstract
In classic instruction following, language like “I’d like the JetBlue flight” maps to actions (e.g., selecting that flight). However, language also conveys information about a user’s underlying reward function (e.g., a general preference for JetBlue), which can allow a model to carry out desirable actions in new contexts. We present a model that infers rewards from language pragmatically: reasoning about how speakers choose utterances not only to elicit desired actions, but also to reveal information about their preferences. On a new interactive flight–booking task with natural language, our model more accurately infers rewards and predicts optimal actions in unseen environments, in comparison to past work that first maps language to actions (instruction following) and then maps actions to rewards (inverse reinforcement learning).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 92 citations
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus et al.ICML 2023 · 36 citations
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths et al.NeurIPS 2022 · 30 citations
- Need Help? Designing Proactive AI Assistants for ProgrammingValerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar et al.CHI 2025 · 23 citations
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 21 citations
Builds on2
Related papers
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan et al.AAAI 2021 · 67 citations
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 7 citations
- Inverse Reinforcement Learning with Natural Language GoalsLi Zhou, Kevin SmallAAAI 2021 · 40 citations
- Using Reinforcement Learning to Train Large Language Models to Explain Human DecisionsJian-Qiao Zhu, Hanbo Xie, Dilip Arumugam, Robert C. Wilson et al.ICLR 2026 · 10 citations
- Causally Robust Reward Learning from Reason-Augmented Preference FeedbackMinjune Hwang, Yigit Korkmaz, Daniel Seita, Erdem BiyikICLR 2026 · 2 citations
