Routing and Reasoned Evaluation with Large Language Models
Guiyao Tie, Tianyao Luo, Xueyang Zhou, Chaoran Hu, Yunhong He, Junran Wu, Yuanfan Yao, Pan Zhou, Lichao Sun
摘要
Large language models (LLMs) are increasingly used to provide automated assessment signals for evaluating model-generated outputs. However, practical deployment faces three persistent challenges: heterogeneous reliability across models, substantial latency and token costs, and the absence of principled strategies for allocating evaluation resources. We introduce REval, a routing-aware automated assessment framework that formulates evaluation as a resource allocation and aggregation problem rather than relying on a single monolithic evaluator. REval combines difficulty-aware routing with reasoned evaluation signals to dynamically select evaluator models on a per-instance basis under explicit accuracy, latency, and cost constraints. Our study makes three contributions. First, we construct six difficulty-aware datasets spanning both reasoning-intensive (mathematics, logic, code) and non-reasoning (knowledge, roleplay, writing) tasks, with human-annotated reference assessments. Second, we provide a systematic empirical analysis of how reasoning traces produced by different evaluator models correlate with assessment outcomes, revealing substantial variance and systematic mismatches across difficulty regimes. Third, we develop and evaluate both offline and online routing strategies that adaptively allocate evaluation queries, achieving substantially improved accuracy–efficiency trade-offs compared to static baselines. Experiments across 19 language models demonstrate that REval significantly reduces evaluation cost and latency while maintaining close alignment with human assessments. These results highlight the importance of routing-aware automated assessment and establish REval as a scalable and reliable framework for large-scale model evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton 等NeurIPS 2025 · 被引用 507 次
- Generative Judge for Evaluating AlignmentJunlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan 等ICLR 2024 · 被引用 173 次
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong 等ICLR 2026 · 被引用 150 次
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual SearchXin Lai, Junyi Li, Wei Li, Tao Liu 等ICLR 2026 · 被引用 124 次
相关 Paper
- Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-JudgeWenbo Zhang, Lijinghua Zhang, Liner Xiang, Hengrui CaiICML 2026 · 被引用 1 次
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMsNigel Fernandez, Branislav Kveton, Ryan A. Rossi, Andrew Lan 等ICLR 2026 · 被引用 6 次
- Think When Needed: Model-Aware Reasoning Routing for LLM-based RankingHuizhong Guo, Tianjun Wei, Dongxia Wang, Yingpeng Du 等SIGIR 2026
- Adaptive Model and Strategy Routing for Cost-Efficient LLM ServicesZhihong Pan, Kai Zhang, Yuze Zhao, Yupeng HanWWW 2026
- Adaptive Prediction-Powered AutoEval with Reliability and Efficiency GuaranteesSangwoo Park, Matteo Zecchin, Osvaldo SimeoneNeurIPS 2025 · 被引用 10 次
