SATER: A Self-Aware and Token-Efficient Approach to Routing and Cascading
Yuanzhe Shen, Yide Liu, Zisu Huang, Ruicheng Yin, Xiaoqing Zheng, Xuanjing Huang
摘要
Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness frequently depends on costly commercial APIs or cloud services. Model selection thus entails a critical trade-off between performance and cost: high-performing LLMs typically incur substantial expenses, whereas budget-friendly small language models (SLMs) are constrained by limited capabilities. Current research primarily proposes two routing strategies: pre-generation routing and cascade routing. Both approaches have distinct characteristics, with cascade routing typically offering superior cost-effectiveness and accuracy despite its higher latency. To further address the limitations of both approaches, we introduce SATER, a dual-mode compatible approach that finetunes models through shortest-response preference optimization and a confidence-aware rejection mechanism. SATER significantly reduces redundant outputs and response times, while improving both the performance of pregeneration routing and the efficiency of cascade routing. Experiments across three SLMs and six datasets, varying in type and complexity, demonstrate that SATER achieves comparable performance while consistently reducing computational costs by over 50% and cascade latency by over 80%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
相关 Paper
- Selective Deferred Routing: Enabling Cost-Efficient Collaboration between Local SLMs and Remote LLMsQijun Miao, Zhixuan FangICML 2026
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal ComputeDujian Ding, Ankur Mallick, Shaokun Zhang, Chi Wang 等ICML 2025
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
- IRT-Router: Effective and Interpretable Multi-LLM Routing via Item Response TheoryWei Song, Zhenya Huang, Cheng Cheng, Weibo Gao 等ACL 2025 · 被引用 20 次
- FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingYujia Fu, Heming Zhong, Dan Huang, Yutong LuACL 2026
