Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
Wenbo Zhang, Lijinghua Zhang, Liner Xiang, Hengrui Cai
Abstract
Reasoning-capable large language models (LLMs) have recently been adopted as automated judges, but their benefits and costs in LLM-as-a-Judge settings remain unclear. Through controlled comparisons between reasoning and non-reasoning judges, we show that explicit reasoning substantially improves judgment accuracy on tasks requiring structured verification (e.g., math and coding), while offering limited or even negative gains on simpler evaluations and incurring significantly higher computational cost . These findings motivate that reasoning should be used selectively rather than universally, with awareness of possible distribution shift . We propose a Robust Adaptive Cost-Efficient Routing (RACER), which dynamically selects between reasoning and non-reasoning judges under a fixed budget by formulating routing as a constrained distributionally robust optimization problem. RACER explicitly accounts for distribution shift via a KL-divergence uncertainty set, admits an efficient primal--dual algorithm, and enjoys theoretical guarantees including uniqueness of the optimal policy and linear convergence. Extensive experiments show that RACER achieves superior accuracy--cost trade-offs under distribution shift.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb17eabd-9ccf-4b3b-b634-ef7cd816f94bBuilds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Prometheus: Inducing Fine-Grained Evaluation Capability in Language ModelsSeungone Kim, Jamin Shin, Yejin Choi, Joel Jang et al.ICLR 2024 · 468 citations
- Large-Scale Methods for Distributionally Robust OptimizationDaniel Levy, Yair Carmon, John C. Duchi, Aaron SidfordNeurIPS 2020 · 281 citations
Related papers
- Routing and Reasoned Evaluation with Large Language ModelsGuiyao Tie, Tianyao Luo, Xueyang Zhou, Chaoran Hu et al.ICML 2026
- RACER: Risk-Aware Calibrated Efficient Routing for Large Language ModelsSai Hao, Hao Zeng, Hongxin Wei, Bingyi JingICML 2026 · 1 citation
- Confidence-Guided Stepwise Model Routing for Cost-Efficient ReasoningSangmook Lee, Dohyung Kim, Hyukhun Koh, Nakyeong Yang et al.AAAI 2026 · 3 citations
- R2-Router: A New Paradigm for LLM Routing with ReasoningJiaqi Xue, Qian Lou, Jiarong Xing, Heng HuangICML 2026 · 12 citations
- Think When Needed: Model-Aware Reasoning Routing for LLM-based RankingHuizhong Guo, Tianjun Wei, Dongxia Wang, Yingpeng Du et al.SIGIR 2026
