Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement
Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Clark
摘要
Our goal is a teachable reasoning system for question-answering (QA), where a user can interact with faithful answer explanations, and correct its errors so that the system improves over time. Our approach is to augment a QA model with a dynamic memory of user feedback, containing user-supplied corrections to erroneous model beliefs that users identify during interaction. Retrievals from memory are used as additional context for QA, to help avoid previous mistakes in similar new situationsa novel application of memory-based continuous learning. With simulated feedback, we find that our system (called TeachMe 1 ) continually improves with time, and without model retraining, requiring feedback on only 25% of training examples to reach within 1% of the upper-bound (feedback on all examples). Similarly, in experiments with real users, we observe a similar trend, with performance improving by over 15% on a hidden test set after teaching. This suggests new opportunities for using frozen language models in an interactive setting where users can inspect, debug, and correct the model's beliefs, leading to improved system's performance over time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri 等NeurIPS 2023 · 被引用 516 次
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 被引用 45 次
- Entailer: Answering Questions with Faithful and Truthful Chains of ReasoningOyvind Tafjord, Bhavana Dalvi Mishra, Peter ClarkEMNLP 2022 · 被引用 28 次
- How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent AdvancesZihan Zhang, Meng Fang, Ling Chen, Mohammad-Reza Namazi-Rad 等EMNLP 2023 · 被引用 22 次
- Faithful Question Answering with Monte-Carlo PlanningRuixin Hong, Hongming Zhang, Hong Zhao, Dong Yu 等ACL 2023 · 被引用 7 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- A Rationale-Centric Framework for Human-in-the-loop Machine LearningJinghui Lu, Linyi Yang, Brian MacNamee, Yue ZhangACL 2022 · 被引用 46 次
相关 Paper
- Continually Improving Extractive QA via Human FeedbackGe Gao, Hung-Ting Chen, Yoav Artzi, Eunsol ChoiEMNLP 2023 · 被引用 5 次
- Memory-assisted prompt editing to improve GPT-3 after deploymentAman Madaan, Niket Tandon, Peter Clark, Yiming YangEMNLP 2022 · 被引用 1 次
- Remembering for the Right Reasons: Explanations Reduce Catastrophic ForgettingSayna Ebrahimi, Suzanne Petryk, Akash Gokul, William Gan 等ICLR 2021 · 被引用 3 次
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language ModelsDeepak Nathani, David Wang, Liangming Pan, William Yang WangEMNLP 2023 · 被引用 6 次
- BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of BeliefNora Kassner, Oyvind Tafjord, Hinrich Schütze, Peter ClarkEMNLP 2021 · 被引用 2 次
