Fantastic Rewards and How to Tame Them: A Case Study on Reward Learning for Task-oriented Dialogue Systems
Yihao Feng, Shentao Yang, Shujian Zhang, Jianguo Zhang, Caiming Xiong, Mingyuan Zhou, Huan Wang
Abstract
When learning task-oriented dialogue (ToD) agents, reinforcement learning (RL) techniques can naturally be utilized to train dialogue strategies to achieve user-specific goals. Prior works mainly focus on adopting advanced RL techniques to train the ToD agents, while the design of the reward function is not well studied. This paper aims at answering the question of how to efficiently learn and leverage a reward function for training end-to-end (E2E) ToD agents. Specifically, we introduce two generalized objectives for reward-function learning, inspired by the classical learning-to-rank literature. Further, we utilize the learned reward function to guide the training of the E2E ToD agent. With the proposed techniques, we achieve competitive results on the E2E response-generation task on the Multiwoz 2.0 dataset. Source code and checkpoints are publicly released at https://github.com/Shentao-YANG/Fantastic_Reward_ICLR2023.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2af78801-544c-487d-a49e-96d1cd6e1016Cited by top-tier papers10
- Plug-and-Play Policy Planner for Large Language Model Powered Dialogue AgentsYang Deng, Wenxuan Zhang, Wai Lam, See-Kiong Ng et al.ICLR 2024 · 86 citations
- A Dense Reward View on Aligning Text-to-Image Diffusion with PreferenceShentao Yang, Tianqi Chen, Mingyuan ZhouICML 2024 · 53 citations
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng et al.NeurIPS 2023 · 39 citations
- Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language ModelsAshutosh Baheti, Ximing Lu, Faeze Brahman, Ronan Le Bras et al.ICLR 2024 · 16 citations
- Planning Like Human: A Dual-process Framework for Dialogue PlanningTao He, Lizi Liao, Yixin Cao, Yuanxing Liu et al.ACL 2024 · 7 citations
Builds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz et al.NeurIPS 2020 · 590 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
- UBAR: Towards Fully End-to-End Task-Oriented Dialog System with GPT-2Yunyi Yang, Yunhao Li, Xiaojun QuanAAAI 2021 · 217 citations
Related papers
- KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningXiao Yu, Qingyang Wu, Kun Qian, Zhou YuEMNLP 2023 · 1 citation
- GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue SystemsYoungsoo Jang, Jongmin Lee, Kee-Eung KimICLR 2022 · 45 citations
- Modelling Hierarchical Structure between Dialogue Policy and Natural Language Generator with Option Framework for Task-oriented Dialogue SystemJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuICLR 2021 · 48 citations
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 12 citations
- Dual-Feedback Knowledge Retrieval for Task-Oriented Dialogue SystemsTianyuan Shi, Liangzhi Li, Zijian Lin, Tao Yang et al.EMNLP 2023 · 9 citations
