Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning
Xuechen Zhang, Zijian Huang, Ege Onur Taga, Carlee Joe-Wong, Samet Oymak, Jiasi Chen
Abstract
Recent successes in natural language processing have led to the proliferation of large language models (LLMs) by multiple providers. Each LLM offering has different inference accuracy, monetary cost, and latency, and their accuracy further depends on the exact wording of the question (i.e., the specific prompt). At the same time, users often have a limit on monetary budget and latency to answer all their questions, and they do not know which LLMs to choose for each question to meet their accuracy and long term budget requirements. To navigate this rich design space, we propose TREACLE (hrifty soning via ontext-Aware LM and Prompt Slection), a reinforcement learning policy that jointly selects the model and prompting scheme while respecting the user's monetary cost and latency constraints. TREACLE uses the problem context, including question text embeddings (reflecting the type or difficulty of a query) and the response history (reflecting the consistency of previous responses) to make smart decisions. Our evaluations on standard reasoning datasets (GSM8K, CSQA, and LLC) with various LLMs and prompts show that TREACLE enables cost savings of up to 85% compared to baselines, while maintaining high accuracy. Importantly, it provides the user with the ability to gracefully trade off accuracy for cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9498f59-80de-4ccb-9d99-b4134fe9c8c2Cited by top-tier papers8
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for ReasoningXuechen Zhang, Zijian Huang, Yingcong Li, Chenshun Ni et al.NeurIPS 2025 · 66 citations
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for ReasoningAntonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang et al.NeurIPS 2025 · 10 citations
- RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMsNigel Fernandez, Branislav Kveton, Ryan A. Rossi, Andrew Lan et al.ICLR 2026 · 6 citations
- QStore: Quantization-Aware Compressed Model StorageRaunak Shah, Zhaoheng Li, Yongjoo ParkVLDB 2026 · 3 citations
- Timely Machine: Awareness of Time Makes Test-Time Scaling AgenticYichuan Ma, Linyang Li, Yongkang Chen, Peiji Li et al.ACL 2026 · 3 citations
Builds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju et al.NeurIPS 2024 · 145 citations
Related papers
- An Empirical Study of LLM Reasoning Ability Under Strict Output Length ConstraintYi Sun, Han Wang, Jiaqiang Li, Jiacheng Liu et al.EMNLP 2025 · 1 citation
- SABER: Switchable and Balanced Training for Efficient LLM ReasoningKai Zhao, Yanjun Zhao, Jiaming Song, Shien He et al.AAAI 2026 · 9 citations
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong et al.ICLR 2026 · 1 citation
- Adaptive Model and Strategy Routing for Cost-Efficient LLM ServicesZhihong Pan, Kai Zhang, Yuze Zhao, Yupeng HanWWW 2026
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
