Towards Boosting the Open-Domain Chatbot with Human Feedback
Hua Lu, Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang
摘要
Many open-domain dialogue models pre-trained with social media comments can generate coherent replies but have difficulties producing engaging responses. This phenomenon might mainly result from the deficiency of annotated human-human conversations and the misalignment with human preference. In this paper, we propose a novel and efficient framework Diamante to boost the open-domain chatbot, where two kinds of human feedback (including explicit demonstration and implicit preference) are collected and leveraged. By asking annotators to select or amend the model-generated candidate responses, Diamante efficiently collects the human demonstrated responses and constructs a Chinese chit-chat dataset. To enhance the alignment with human preference, Diamante leverages the implicit preference in the data collection process and introduces the generation-evaluation joint training. Comprehensive experiments indicate that the Diamante dataset and joint training paradigm can significantly boost the performance of pre-trained dialogue models. The overall engagingness of the previous state-of-the-art model has been improved remarkably by 50% in Chinese open-domain conversations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesHannah Kirk, Andrew M. Bean, Bertie Vidgen, Paul Röttger 等EMNLP 2023 · 被引用 13 次
- IEvoAgent: Evolving Conversational Agent based on User Implicit FeedbackYichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo 等ACL 2026
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston 等ACL 2020 · 被引用 18 次
- Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedbackJing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora 等ACL 2023 · 被引用 13 次
相关 Paper
- DialogVED: A Pre-trained Latent Variable Encoder-Decoder Model for Dialog Response GenerationWei Chen, Yeyun Gong, Song Wang, Bolun Yao 等ACL 2022
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- ProphetChat: Enhancing Dialogue Generation with Simulation of Future ConversationChang Liu, Xu Tan, Chongyang Tao, Zhenxin Fu 等ACL 2022
- Dialogue Response Ranking Training with Large-Scale Human Feedback DataXiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett 等EMNLP 2020 · 被引用 67 次
- Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from ScratchXueru Wen, Jie Lou, Zichao Li, Yaojie Lu 等ACL 2025 · 被引用 1 次
