The Avengers: A Routing Recipe for Collective Intelligence in Language Models
Yiqun Zhang, Hao Li, Chenxu Wang, Linyao Chen, Qiaosheng Zhang, Peng Ye, Shi Feng, Xinrun Wang, Xu Jia, Lei Bai, Shuyue Hu
Abstract
Proprietary models are increasingly dominating the race for ever-larger language models. Can open-source, smaller models remain competitive across a broad range of tasks? In this paper, we present the Avengers---a lightweight framework that leverages the collective intelligence of these smaller models. The Avengers builds upon four lightweight operations: (i) embedding: encode queries using a text embedding model; (ii) clustering: group queries based on their semantic similarity; (iii) scoring: scores each model's performance within each cluster; and (iv) voting: improve outputs via repeated sampling and voting. At inference time, each query is embedded and assigned to its nearest cluster. The top-performing model(s) within that cluster are selected to generate the response with repeated sampling. Remarkably, with 10 open-source models ( 7B parameters each), the Avengers surpasses GPT-4o, 4.1, and 4.5 in average performance across 15 diverse datasets spanning mathematics, coding, logical reasoning, general knowledge, and affective tasks. In particular, it surpasses GPT-4.1 on mathematics tasks by 18.21% and on code tasks by 7.46%. Furthermore, the Avengers delivers superior out-of-distribution generalization, and remains robust across various embedding models, clustering algorithms, ensemble strategies, data efficiency, and values of its sole parameter---the number of clusters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b96bd424-b843-4ecf-8cf3-838891da6214Cited by top-tier papers1
Ask how each one uses itBuilds on17
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok et al.NeurIPS 2024 · 113 citations
- Universal Model Routing for Efficient LLM InferenceWittawat Jitkrittum, Harikrishna Narasimhan, Ankit Singh Rawat, Jeevesh Juneja et al.ICLR 2026 · 99 citations
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 95 citations
Related papers
- EffGen: Enabling Small Language Models as Capable Autonomous AgentsGaurav Srivastava, Aafiya Hussain, Chi Wang, Yingyan (Celine) Lin et al.ICML 2026
- Agent Lumos: Unified and Modular Training for Open-Source Language AgentsDa Yin, Faeze Brahman, Abhilasha Ravichander, Khyathi Raghavi Chandu et al.ACL 2024
- Mixture-of-Agents Enhances Large Language Model CapabilitiesJunlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang et al.ICLR 2025
- PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem SolvingMihir Parmar, Palash Goyal, Xin Liu, Yiwen Song et al.EMNLP 2025
- Beyond Gemini-3-Pro: Revisiting LLM Routing and Aggregation at ScaleShengji Tang, Weihao Lin, Peng Ye, Jingqi Ye et al.ICML 2026
