Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback
Jing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora, Y-Lan Boureau, Jason Weston
Abstract
Frozen models trained to mimic static datasets can never improve their performance. Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance. In this work we study how to improve internet-driven conversational skills in such a learning framework. We collect deployment data, which we make publicly available, of human interactions, and collect various types of human feedback – including binary quality measurements, free-form text feedback, and fine-grained reasons for failure. We then study various algorithms for improving from such feedback, including standard supervised learning, rejection sampling, model-guiding and reward-based learning, in order to make recommendations on which type of feed- back and algorithms work best. We find the recently introduced DIRECTOR model (Arora et al., 2022) shows significant improvements over other existing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49f6a704-6a3b-4ce8-844d-c7c455b1a5b2Cited by top-tier papers13
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 303 citations
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsXuan Zhang, Chao Du, Tianyu Pang, Qian Liu et al.NeurIPS 2024 · 177 citations
- Aligning LLM Agents by Learning Latent Preference from User EditsGe Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro et al.NeurIPS 2024 · 102 citations
- I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-ImitationChandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras et al.ACL 2023 · 18 citations
Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 329 citations
- Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend SkillsEric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston et al.ACL 2020 · 18 citations
- I like fish, especially dolphins: Addressing Contradictions in Dialogue ModelingYixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela et al.ACL 2021
Related papers
- Internet-Augmented Dialogue GenerationMojtaba Komeili, Kurt Shuster, Jason WestonACL 2022
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 5 citations
- Iterative Label Refinement Matters More than Preference Optimization under Weak SupervisionYaowen Ye, Cassidy Laidlaw, Jacob SteinhardtICLR 2025
- Continual Learning for Instruction Following from Realtime FeedbackAlane Suhr, Yoav ArtziNeurIPS 2023 · 27 citations
- Continual Dialogue State Tracking via Example-Guided Question AnsweringHyundong Cho, Andrea Madotto, Zhaojiang Lin, Khyathi Raghavi Chandu et al.EMNLP 2023
