Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
Sanjiban Choudhury, Paloma Sodhi
摘要
While large language models (LLMs) show impressive decision-making abilities, current methods lack a mechanism for automatic self-improvement from errors during task execution. We propose LEAP, an iterative fine-tuning framework that continually improves LLM agents using feedback from AI expert teachers. Our key insight is to equip the expert teachers with a privileged state -- information that is available during training but hidden at test time. This allows even weak experts to provide precise guidance, significantly improving the student agent's performance without access to privileged information at test time. We evaluate LEAP on diverse decision-making benchmarks, including text-based games (ALFWorld), web navigation (WebShop), and interactive coding (Intercode Bash). Our experiments show that LEAP (1) outperforms behavior cloning and ReAct baselines (2) enables weak student models (e.g., Llama3-8B) to exceed the performance of strong teacher models (GPT4-o), and (3) allows weak models to self-improve using privileged versions of themselves. We also provide a theoretical analysis showing that LEAP's success hinges on balancing privileged information with the student's realizability, which we empirically validate. Our code is available at https://leap-llm.github.io
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-TuningGokul Swamy, Sanjiban Choudhury, Wen Sun, Steven Wu 等ICLR 2026 · 被引用 66 次
- Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy OptimizationZeyuan Liu, Jeonghye Kim, Xufang Luo, Dongsheng Li 等ICLR 2026 · 被引用 18 次
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge GraphsYurun Chen, Xueyu Hu, Yuhan Liu, Ziqi Wang 等CVPR 2026 · 被引用 7 次
- Preemptive Detection and Correction of Misaligned Actions in LLM AgentsHaishuo Fang, Xiaodan Zhu, Iryna GurevychEMNLP 2025 · 被引用 1 次
- Multi-Turn Code Generation Through Single-Step RewardsArnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen, Alexander M. Rush 等ICML 2025
它引用的顶会 Paper24
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
相关 Paper
- Policy Improvement using Language Feedback ModelsVictor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre CôtéNeurIPS 2024 · 被引用 18 次
- The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided ImprovementRuihan Yang, Fanghua Ye, Jian Li, Siyu Yuan 等NeurIPS 2025 · 被引用 21 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- A Critical Evaluation of AI Feedback for Aligning Large Language ModelsArchit Sharma, Sedrick Scott Keh, Eric Mitchell, Chelsea Finn 等NeurIPS 2024 · 被引用 50 次
- AgentRefine: Enhancing Agent Generalization through Refinement TuningDayuan Fu, Keqing He, Yejie Wang, Wentao Hong 等ICLR 2025
