Learning to Decode Collaboratively with Multiple Language Models
Zejiang Shen, Hunter Lang, Bailin Wang, Yoon Kim, David A. Sontag
Abstract
We propose a method to teach multiple large language models (LLM) to collaborate by interleaving their generations at the token level. We model the decision of which LLM generates the next token as a latent variable. By optimizing the marginal likelihood of a training set under our latent variable model, the base LLM automatically learns when to generate itself and when to call on one of the "assistant" language models to generate, all without direct supervision. Token-level collaboration during decoding allows for a fusion of each model's expertise in a manner tailored to the specific task at hand. Our collaborative decoding is especially useful in cross-domain settings where a generalist base LLM learns to invoke domain expert models. On instruction-following, domainspecific QA, and reasoning tasks, we show that the performance of the joint system exceeds that of the individual models. Through qualitative analysis of the learned latent decisions, we show models trained with our method exhibit several interesting collaboration patterns, e.g., template-filling. 1 Here's a recipe for Kubdari, a traditional Georgian dish: Ingredients: * 1 lb ground beef * 1 onion, finely cho pped * 2 cloves garlic, minced * 1 cup chopped parsley * 1 cup chopped cilantro * 1 cup chopped dill * 1 cup chopped ...
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc89da83-d1c4-4eda-8296-c57ab03e4e33Cited by top-tier papers25
- Cache-to-Cache: Direct Semantic Communication Between Large Language ModelsTianyu Fu, Zihan Min, Hanling Zhang, Jichao Yan et al.ICLR 2026 · 59 citations
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- Smoothie: Label Free Language Model RoutingNeel Guha, Mayee F. Chen, Trevor Chow, Ishan S. Khare et al.NeurIPS 2024 · 44 citations
- When One LLM Drools, Multi-LLM Collaboration RulesShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang et al.ACL 2026 · 27 citations
- Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsShangbin Feng, Zifeng Wang, Palash Goyal, Yike Wang et al.NeurIPS 2025 · 26 citations
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Token-Level LLM Collaboration via FusionRouteNuoya Xiong, Yuhang Zhou, Hanqing Zeng, Zhaorun Chen et al.ICML 2026 · 7 citations
- Speculate, then Collaborate: Fusing Knowledge of Language Models during DecodingZiyao Wang, Muneeza Azmat, Ang Li, Raya Horesh et al.ICML 2025
- Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language ModelsJerry Huang, Prasanna Parthasarathi, Mehdi Rezagholizadeh, Sarath ChandarEMNLP 2024 · 6 citations
- Synergistic Weak-Strong Collaboration by Aligning PreferencesYizhu Jiao, Xuchao Zhang, Zhaoyang Wang, Yubo Ma et al.ACL 2025
- Don't Throw Away Your Pretrained ModelShangbin Feng, Wenhao Yu, Yike Wang, Hongming Zhang et al.ICLR 2026 · 10 citations
