Inferring Rewards from Language in Context
Jessy Lin, Daniel Fried, Dan Klein, Anca D. Dragan
摘要
In classic instruction following, language like “I’d like the JetBlue flight” maps to actions (e.g., selecting that flight). However, language also conveys information about a user’s underlying reward function (e.g., a general preference for JetBlue), which can allow a model to carry out desirable actions in new contexts. We present a model that infers rewards from language pragmatically: reasoning about how speakers choose utterances not only to elicit desired actions, but also to reveal information about their preferences. On a new interactive flight–booking task with natural language, our model more accurately infers rewards and predicts optimal actions in unseen environments, in comparison to past work that first maps language to actions (instruction following) and then maps actions to rewards (inverse reinforcement learning).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Inverse Preference Learning: Preference-based RL without a Reward FunctionJoey Hejna, Dorsa SadighNeurIPS 2023 · 被引用 92 次
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- Need Help? Designing Proactive AI Assistants for ProgrammingValerie Chen, Alan Zhu, Sebastian Zhao, Hussein Mozannar 等CHI 2025 · 被引用 23 次
- Reward Design with Language ModelsMinae Kwon, Sang Michael Xie, Kalesha Bullard, Dorsa SadighICLR 2023 · 被引用 21 次
它引用的顶会 Paper2
相关 Paper
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan 等AAAI 2021 · 被引用 67 次
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 被引用 7 次
- Inverse Reinforcement Learning with Natural Language GoalsLi Zhou, Kevin SmallAAAI 2021 · 被引用 40 次
- Using Reinforcement Learning to Train Large Language Models to Explain Human DecisionsJian-Qiao Zhu, Hanbo Xie, Dilip Arumugam, Robert C. Wilson 等ICLR 2026 · 被引用 10 次
- Causally Robust Reward Learning from Reason-Augmented Preference FeedbackMinjune Hwang, Yigit Korkmaz, Daniel Seita, Erdem BiyikICLR 2026 · 被引用 2 次
