Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMs
Runyu Peng, Yunhua Zhou, Kai Lv, Yang Gao, Qipeng Guo, Xipeng Qiu
Abstract
The rapid advancement of Large Language Models (LLMs) has significantly enhanced performance across various natural language processing (NLP) tasks, yet the high computational costs and latency associated with deploying such models continue to pose critical bottlenecks, limiting their broader applicability. To mitigate these challenges, we propose a dynamic hybrid inference framework, Firewall Routing, which efficiently selects between a strong and a weak LLMs based on the complexity of the query. A lightweight routing model is trained to optimize resource allocation by learning from response quality and preventing longtail queries, which are often too hard to solve by LLMs, from being routed to the stronger model. Moreover, our method incorporates multiple sampling to enhance query evaluation reliability while leveraging Hard Blocking and Soft Blocking to handle long-tail queries along with refining labels for model selection. Extensive experiments show our method outperforms existing routing strategies by up to 5.29% in APGR, demonstrating state-of-the-art performance across multiple benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a00f29db-6d4f-4413-95dc-8d7b858be630Builds on1
Related papers
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim et al.ICLR 2024 · 282 citations
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang et al.ICML 2025
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang et al.ICLR 2025
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response TheoryWei Song, Zhenya Huang, Cheng Cheng, Weibo Gao et al.ACL 2025 · 20 citations
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
