Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
Selim Furkan Tekin, Gaowen Liu, Ramana Kompella, Ling Liu
摘要
The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces RL-Focal, a two-stage RL agent framework that routes and ensembles LLMs. First, we develop the Decider RL-agent, which learns to dynamically select an ensemble of small size (m i ) among N LLMs (m i ≪ N ) for incoming queries from a user-defined downstream task i, by maximizing both error-diversity and reasoning-performance of the selected ensemble through iterative updates of task-adaptive rewards and policy. Second, to enable effective fusion of dynamically selected LLMs, we develop the stage-2 Fusion RL-agent, which learns to resolve reasoning conflicts from different LLMs and dynamically adapt to different ensemble teams composed by the Decider Agent for different downstream tasks. Third, we introduce the focal diversity metric to better model the error correlations among multiple LLMs further improving the generalization performance of the Decider Agent, which actively prunes the ensemble combinations. By focal diversity, we enhance performance across tasks by effectively promoting reward-aware and policy-adaptive ensemble selection and inference fusion. Extensive evaluations on five benchmarks show that RL-Focal achieves the performance improvement of 8.48% with an ensemble of small size compared to the best individual LLM in a pool and offers stronger robustness. Code is available at https://github.com/sftekin/rl-focal Ensemble Learning in LLMs: Related Work and Open Challenges. Most popular ensemble training methods are the distillation (the ensemble at each token generation step) (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
相关 Paper
- RLAE: Reinforcement Learning-Assisted Ensemble for LLMsYuqian Fu, Yuanheng Zhu, Jiajun Chai, Guojun Yin 等EMNLP 2025 · 被引用 1 次
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement LearningHaozhen Zhang, Tao Feng, Jiaxuan YouNeurIPS 2025 · 被引用 81 次
- Demonstration Selection for In-Context Learning via Reinforcement LearningXubin Wang, Jianfei Wu, Yichen Yuan, Deyu Cai 等ICML 2025
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang 等NeurIPS 2024 · 被引用 94 次
- Balancing Act: Diversity and Consistency in Large Language Model EnsemblesAhmed Abdulaal, Chen Jin, Nina Montaña Brown, Aryo Pradipta Gema 等ICLR 2025
