Human-centric dialog training via offline reinforcement learning
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, Rosalind W. Picard
摘要
How can we train a dialog model to produce better conversations by learning from human feedback, without the risk of humans teaching it harmful chat behaviors? We start by hosting models online, and gather human feedback from real-time, open-ended conversations, which we then use to train and improve the models using offline reinforcement learning (RL). We identify implicit conversational cues including language similarity, elicitation of laughter, sentiment, and more, which indicate positive human feedback, and embed these in multiple reward functions. A wellknown challenge is that learning an RL policy in an offline setting usually fails due to the lack of ability to explore and the tendency to make over-optimistic estimates of future reward. These problems become even harder when using RL for language models, which can easily have a 20,000 action vocabulary and many possible reward functions. We solve the challenge by developing a novel class of offline RL algorithms. These algorithms use KL-control to penalize divergence from a pretrained prior language model, and use a new strategy to make the algorithm pessimistic, instead of optimistic, in the face of uncertainty. We test the resulting dialog model with ratings from 80 users in an open-domain setting and find it achieves significant improvements over existing deep offline RL approaches. The novel offline RL method is viable for improving any existing generative dialog model using a static dataset of human feedback.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
- ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RLYifei Zhou, Andrea Zanette, Jiayi Pan, Sergey Levine 等ICML 2024 · 被引用 163 次
- Generalized Decision Transformer for Offline Hindsight Information MatchingHiroki Furuta, Yutaka Matsuo, Shixiang Shane GuICLR 2022 · 被引用 125 次
它引用的顶会 Paper2
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RLSeyed Kamyar Seyed Ghasemipour, Dale Schuurmans, Shixiang Shane GuICML 2021 · 被引用 138 次
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
相关 Paper
- Offline RL for Natural Language Generation with Implicit Language Q LearningCharlie Snell, Ilya Kostrikov, Yi Su, Sherry Yang 等ICLR 2023 · 被引用 9 次
- GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue SystemsYoungsoo Jang, Jongmin Lee, Kee-Eung KimICLR 2022 · 被引用 45 次
- WildReward: Learning Reward Models from In-the-Wild Human InteractionsHao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao 等ACL 2026 · 被引用 3 次
- On the Effectiveness of Offline RL for Dialogue Response GenerationPaloma Sodhi, Felix Wu, Ethan R. Elenberg, Kilian Q. Weinberger 等ICML 2023 · 被引用 6 次
- Building Persona Consistent Dialogue Agents with Offline Reinforcement LearningRyan Shea, Zhou YuEMNLP 2023 · 被引用 4 次
