MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning
Nian Ran, Zhongzheng Li, Yue Wang, Qingsong Ran, Xiaoyuan Zhang, Shikun Feng, Richard Allmendinger, Xiaoguang Zhao
Abstract
Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured combinatorial spaces. Traditional evolutionary algorithms often get trapped in local optima, while expert knowledge can provide crucial guidance for accelerating convergence. Large language models (LLMs) offer powerful priors and reasoning ability, making them natural optimizers when expert knowledge matters. However, closed-source LLMs, though strong in exploration, cannot update their parameters and thus cannot internalize experience. Conversely, smaller open models can be continually fine-tuned but lack broad knowledge and reasoning strength. We introduce Multi-LLM Collaborative Co-evolution (MCCE), a hybrid framework that unites a frozen closed-source LLM with a lightweight trainable model. The system maintains a trajectory memory of past search processes; the small model is progressively refined via reinforcement learning, with the two models jointly supporting and complementing each other in global exploration. Unlike model distillation, this process enhances the capabilities of both models through mutual inspiration. Experiments on multi-objective drug design benchmarks show that MCCE achieves state-of-the-art Pareto front quality and consistently outperforms baselines. These results highlight a new paradigm for enabling continual evolution in hybrid LLM systems, combining knowledge-driven exploration with experience-driven learning. The code of MCCE is available on https://github.com/lzz-z/MCCE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e6568bd-05a1-400b-967b-119fd47ae8feBuilds on17
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang et al.NeurIPS 2025 · 310 citations
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language ModelFei Liu, Xialiang Tong, Mingxuan Yuan, Xi Lin et al.ICML 2024 · 238 citations
- Pareto Set Learning for Neural Multi-Objective Combinatorial OptimizationXi Lin, Zhiyuan Yang, Qingfu ZhangICLR 2022 · 105 citations
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsLu Ma, Hao Liang, Meiyi Qiang, Lexiang Tang et al.ICLR 2026 · 103 citations
Related papers
- Efficient Evolutionary Search Over Chemical Space with Large Language ModelsHaorui Wang, Marta Skreta, Cher Tian Ser, Wenhao Gao et al.ICLR 2025
- MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide DesignGen Zhou, Sugitha Janarthanan, Lianghong Chen, Pingzhao HuICLR 2026 · 1 citation
- CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic DesignZiyao Huang, Weiwei Wu, Kui Wu, Wei-Bin Lee et al.ICLR 2026 · 41 citations
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific ReasoningKehua Feng, Keyan Ding, Zhihui Zhu, Lei Liang et al.ICLR 2026 · 4 citations
- Self-Evolutionary Reinforced Knowledge Distillation for Multi-Modal Tool-Use AgentsLei Shen, Chengyu Wang, Yuanjie Lyu, Yuanhao Yue et al.KDD 2026
