Self-Evolutionary Reinforced Knowledge Distillation for Multi-Modal Tool-Use Agents
Lei Shen, Chengyu Wang, Yuanjie Lyu, Yuanhao Yue, Jun Huang, Zhongmin Cai
摘要
Autonomous multi-modal agents are increasingly important in real-world applications due to their ability to reason about complex environments and orchestrate tool use. However, deploying multi-modal large language models (MLLMs) for tool use is often constrained by computational cost and inference latency, creating a pressing need for compact models that retain strong agentic capabilities. Training small multi-modal agents remains difficult: limited backbone capacity weakens multi-step reasoning, reward signals for tool use are often sparse and brittle, and naive distillation can fail to transfer the procedural knowledge required for reliable tool invocation and grounding. In this paper, we propose a two-stage self-evolutionary knowledge distillation framework that equips small MLLMs with robust and adaptive tool-use behaviors. Our method combines (i) mutual information-guided trajectory distillation, which selectively transfers high-utility segments of agentic trajectories from a larger teacher, and (ii) reinforcement-driven policy evolution with iterative teacher feedback. To stabilize learning and prevent semantic collapse, we introduce weighted semantic objectives and iteratively expand competence through error-driven optimization, hybrid experience replay, and group-relative policy refinement with multi-dimensional rewards over answer correctness, invocation validity, and tool effectiveness. Integrated with interactive tool modules, our approach enables small models to achieve strong performance across diverse tool-use benchmarks. Comprehensive experiments show consistent improvements over single-pass distillation and RL baselines. Overall, our framework provides a practical path to deploy efficient multi-modal agents without sacrificing tool-use reliability.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Distilling LLM Agent into Small Models with Retrieval and Code ToolsMinki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho 等NeurIPS 2025 · 被引用 51 次
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He 等ICCV 2025 · 被引用 9 次
- From Interactions to Principles: Experience-Driven Self-Distillation for Evolving LLM AgentsRong Wu, Xiaoman Wang, Jianbiao Mei, Pinlong Cai 等ICML 2026
- Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMsYuanjie Lyu, Chengyu Wang, Jun Huang, Tong XuICML 2026 · 被引用 9 次
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
