Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
Xuechen Zhang, Zijian Huang, Ege Onur Taga, Carlee Joe-Wong, Samet Oymak, Jiasi Chen
摘要
Recent successes in natural language processing have led to the proliferation of large language models (LLMs) by multiple providers. Each LLM offering has different inference accuracy, monetary cost, and latency, and their accuracy further depends on the exact wording of the question (i.e., the specific prompt). At the same time, users often have a limit on monetary budget and latency to answer all their questions, and they do not know which LLMs to choose for each question to meet their accuracy and long term budget requirements. To navigate this rich design space, we propose TREACLE (hrifty soning via ontext-Aware LM and Prompt Slection), a reinforcement learning policy that jointly selects the model and prompting scheme while respecting the user's monetary cost and latency constraints. TREACLE uses the problem context, including question text embeddings (reflecting the type or difficulty of a query) and the response history (reflecting the consistency of previous responses) to make smart decisions. Our evaluations on standard reasoning datasets (GSM8K, CSQA, and LLC) with various LLMs and prompts show that TREACLE enables cost savings of up to 85% compared to baselines, while maintaining high accuracy. Importantly, it provides the user with the ability to gracefully trade off accuracy for cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for ReasoningXuechen Zhang, Zijian Huang, Yingcong Li, Chenshun Ni 等NeurIPS 2025 · 被引用 66 次
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for ReasoningAntonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang 等NeurIPS 2025 · 被引用 10 次
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMsNigel Fernandez, Branislav Kveton, Ryan A. Rossi, Andrew Lan 等ICLR 2026 · 被引用 6 次
- QStore: Quantization-Aware Compressed Model StorageRaunak Shah, Zhaoheng Li, Yongjoo ParkVLDB 2026 · 被引用 3 次
- Timely Machine: Awareness of Time Makes Test-Time Scaling AgenticYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju 等NeurIPS 2024 · 被引用 145 次
相关 Paper
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length ConstraintYi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu 等EMNLP 2025 · 被引用 1 次
- SABER: Switchable and Balanced Training for Efficient LLM ReasoningKai Zhao, Yanjun Zhao, Jiaming Song, Shien He 等AAAI 2026 · 被引用 9 次
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong 等ICLR 2026 · 被引用 1 次
- Adaptive Model and Strategy Routing for Cost-Efficient LLM ServicesZhihong Pan, Kai Zhang, Yuze Zhao, Yupeng HanWWW 2026
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
