Learning from Mistakes via Cooperative Study Assistant for Large Language Models
Danqing Wang, Lei Li
摘要
Large language models (LLMs) have demonstrated their potential to refine their generation based on their own feedback. However, the feedback from LLM itself is often inaccurate, thereby limiting its benefits. In this paper, we propose Study Assistant for Large LAnguage Model (SALAM), a novel framework with an auxiliary agent to assist the main LLM in learning from mistakes through interactive cooperation. In the gathering phase, the student assistant agent probes the main LLM, analyzes its errors, and collects the interaction in a mistake memory. During the examination phase, the study assistant provides guidelines by retrieving relevant cases to help the main LLM anticipate and avoid similar errors. We first investigate the effectiveness of a general study assistant and then customize it to provide LLMspecific guidance through imitation learning from successful guidance experiences. Our experiments on three LLMs using two challenging frameworks demonstrate that SALAM can significantly boost LLMs by an accuracy margin of up to 6.6 on BBH and 12.6 on BBQ 1 . 1 https://dqwang122.github.io/projects/SALAM . Jane thought today is 3/11/2002, but today is in fact Mar 12, which is 1 day later. What is the date a month ago? 02/11/2002 False Guideline: For dates in a problem, identify the correct date from which calculations should be made.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi 等EMNLP 2024 · 被引用 119 次
- Temporal Knowledge Question Answering via Abstract Reasoning InductionZiyang Chen, Dongfang Li, Xiang Zhao, Baotian Hu 等ACL 2024 · 被引用 9 次
- Mitigating Social Bias in Large Language Models: A Multi-Objective Approach Within a Multi-Agent FrameworkZhenjie Xu, Wenqing Chen, Yi Tang, Xuanying Li 等AAAI 2025 · 被引用 5 次
- QueryAgent: A Reliable and Efficient Reasoning Framework with Environmental Feedback based Self-CorrectionXiang Huang, Sitao Cheng, Shanshan Huang, Jiayu Shen 等ACL 2024 · 被引用 3 次
- Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale TuningSohan Patnaik, Milan Aggarwal, Sumit Bhatia, Balaji KrishnamurthyACL 2025 · 被引用 2 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
相关 Paper
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 被引用 95 次
- Generator-Assistant Stepwise Rollback Framework for Large Language Model AgentXingzuo Li, Kehai Chen, Yunfei Long, Xuefeng Bai 等EMNLP 2025 · 被引用 1 次
- Boosting the Potential of Large Language Models with an Intelligent Information AssistantYujia Zhou, Zheng Liu, Zhicheng DouNeurIPS 2024
- Experiential Co-Learning of Software-Developing AgentsChen Qian, Yufan Dang, Jiahao Li, Wei Liu 等ACL 2024
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based AgentsTao Wu, Jingyuan Chen, Wang Lin, Mengze Li 等ACL 2025 · 被引用 16 次
