Learning Rewards From Linguistic Feedback
Theodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan, Thomas L. Griffiths
摘要
We explore unconstrained natural language feedback as a learning signal for artificial agents. Humans use rich and varied language to teach, yet most prior work on interactive learning from language assumes a particular form of input (e.g., commands). We propose a general framework which does not make this assumption, instead using aspect-based sentiment analysis to decompose feedback into sentiment over the features of a Markov decision process. We then infer the teacher's reward function by regressing the sentiment on the features, an analogue of inverse reinforcement learning. To evaluate our approach, we first collect a corpus of teaching behavior in a cooperative task where both teacher and learner are human. We implement three artificial learners: sentiment-based "literal" and "pragmatic" models, and an inference network trained end-to-end to predict rewards. We then re-run our initial experiment, pairing human teachers with these artificial learners. All three models successfully learn from interactive human feedback. The inference network approaches the performance of the "literal" sentiment model, while the "pragmatic" model nears human performance. Our work provides insight into the information structure of naturalistic linguistic feedback as well as methods to leverage it for reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus 等ICML 2023 · 被引用 36 次
- Interactive Learning from Activity DescriptionKhanh Nguyen, Dipendra Misra, Robert E. Schapire, Miroslav Dudík 等ICML 2021 · 被引用 36 次
- How to talk so AI will learn: Instructions, descriptions, and autonomyTheodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Tom Griffiths 等NeurIPS 2022 · 被引用 30 次
- A Framework for Learning to Request Rich and Contextually Useful Information from HumansKhanh X. Nguyen, Yonatan Bisk, Hal Daumé IIIICML 2022 · 被引用 21 次
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 等ICLR 2024 · 被引用 20 次
它引用的顶会 Paper2
相关 Paper
- Inferring Rewards from Language in ContextJessy Lin, Daniel Fried, Dan Klein, Anca D. DraganACL 2022 · 被引用 71 次
- Multi-Agent Learning from LearnersMine Melodi Caliskan, Francesco Chini, Setareh MaghsudiICML 2023
- Inverse Optimal Control Adapted to the Noise Characteristics of the Human Sensorimotor SystemMatthias Schultheis, Dominik Straub, Constantin A. RothkopfNeurIPS 2021 · 被引用 25 次
- Retrospective Learning from InteractionsZizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng 等ACL 2025
- Maximum Likelihood Constraint Inference for Inverse Reinforcement LearningDexter R. R. Scobee, S. Shankar SastryICLR 2020 · 被引用 74 次
