An Architecture Search Framework for Inference-Time Techniques
Jon Saad-Falcon, Adrian Gamarra Lafuente, Shlok Natarajan, Nahum Maru, Hristo Todorov, Etash Kumar Guha, Estefany Kelly Buchanan, Mayee F. Chen, Neel Guha, Christopher Ré, Azalia Mirhoseini
Abstract
Inference-time techniques, such as repeated sampling or iterative revisions, are emerging as powerful ways to enhance large-language models (LLMs) at test time. However, best practices for developing systems that combine these techniques remain underdeveloped due to our limited understanding of the utility of each technique across models and tasks, the interactions between them, and the massive search space for combining them. To address these challenges, we introduce ARCHON, a modular and automated framework for optimizing the process of selecting and combining inference-time techniques and LLMs. Given a compute budget and a set of available LLMs, ARCHON explores a large design space to discover optimized configurations tailored to target benchmarks. It can design custom or general-purpose architectures that advance the Pareto frontier of accuracy vs. maximum token budget compared to top-performing baselines. Across instructionfollowing, reasoning, and coding tasks, we show that ARCHON can leverage additional inference compute budget to design systems that outperform frontier models such as OpenAI's o1, GPT-4o, and Claude 3.5 Sonnet by an average of 15.1%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Multi-Agent Design: Optimizing Agents with Better Prompts and TopologiesHan Zhou, Xingchen Wan, Ruoxi Sun, Hamid Palangi et al.ICLR 2026 · 127 citations
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making TasksVishnu Sarukkai, Zhiqiang Xie, Kayvon FatahalianNeurIPS 2025 · 22 citations
- AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex TasksFali Wang, Hui Liu, Zhenwei Dai, Jingying Zeng et al.NeurIPS 2025 · 20 citations
- Abacus: A Cost-Based Optimizer for Semantic Operator SystemsMatthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano et al.VLDB 2026 · 9 citations
- An Information Theoretic Perspective on Agentic System DesignShizhe He, Avanika Narayan, Ishan S. Khare, Scott W. Linderman et al.ICLR 2026 · 6 citations
Builds on4
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- GAIA: a benchmark for General AI AssistantsGrégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun et al.ICLR 2024 · 716 citations
- Mixture-of-Agents Enhances Large Language Model CapabilitiesJunlin Wang, Jue Wang, Ben Athiwaratkun, Ce Zhang et al.ICLR 2025
- Automated Design of Agentic SystemsShengran Hu, Cong Lu, Jeff CluneICLR 2025
Related papers
- When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMsAmmar Khairi, Daniel D'souza, Ye Shen, Julia Kreutzer et al.EMNLP 2025 · 1 citation
- OptScale: Probabilistic Optimality for Inference-time ScalingYoukang Wang, Jian Wang, Rubing Chen, Xiao-Yong WeiAAAI 2026 · 2 citations
- Inference Scaling for Long-Context Retrieval Augmented GenerationZhenrui Yue, Honglei Zhuang, Aijun Bai, Kai Hui et al.ICLR 2025
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-SolvingYangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck et al.ICLR 2025
- LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance ModelingXin Wang, Zhenhao Li, Zishuo DingICSE 2026
