C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
Antonios Valkanas, Soumyasundar Pal, Pavel Rumiantsev, Yingxue Zhang, Mark Coates
摘要
Large language models (LLMs) have achieved impressive results on complex reasoning tasks, but their high inference cost remains a major barrier to real-world deployment. A promising solution is to use cascaded inference, where small, cheap models handle easy queries, and only the hardest examples are escalated to more powerful models. However, existing cascade methods typically rely on supervised training with labeled data, offer no theoretical generalization guarantees, and provide limited control over test-time computational cost. We introduce C3PO (Cost Controlled Cascaded Prediction Optimization), a self-supervised framework for optimizing LLM cascades under probabilistic cost constraints. By focusing on minimizing regret with respect to the most powerful model (MPM), C3PO avoids the need for labeled data by constructing a cascade using only unlabeled model outputs. It leverages conformal prediction to bound the probability that inference cost exceeds a user-specified budget. We provide theoretical guarantees on both cost control and generalization error, and show that our optimization procedure is effective even with small calibration sets. Empirically, C3PO achieves state-of-the-art performance across a diverse set of reasoning benchmarks including GSM8K, MATH-500, BigBench-Hard and AIME, outperforming strong LLM cascading baselines in both accuracy and cost-efficiency. Our results demonstrate that principled, label-free cascade optimization can enable scalable LLM deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Hybrid LLM: Cost-Efficient and Quality-Aware Query RoutingDujian Ding, Ankur Mallick, Chi Wang, Robert Sim 等ICLR 2024 · 被引用 282 次
- C3oT: Generating Shorter Chain-of-Thought Without Compromising EffectivenessYu Kang, Xianghui Sun, Liangyu Chen, Wei ZouAAAI 2025 · 被引用 162 次
- Large Language Model Cascades with Mixture of Thought Representations for Cost-Efficient ReasoningMurong Yue, Jie Zhao, Min Zhang, Liang Du 等ICLR 2024 · 被引用 153 次
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju 等NeurIPS 2024 · 被引用 145 次
相关 Paper
- Conformal Constrained Policy Optimization for Cost-Effective LLM AgentsWenwen Si, Sooyong Jang, Insup Lee, Osbert BastaniAAAI 2026 · 被引用 3 次
- Online Cascade Learning for Efficient Inference over StreamsLunyiu Nie, Zhimin Ding, Erdong Hu, Christopher M. Jermaine 等ICML 2024 · 被引用 20 次
- OptScale: Probabilistic Optimality for Inference-time ScalingYoukang Wang, Jian Wang, Rubing Chen, Xiao-Yong WeiAAAI 2026 · 被引用 2 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning TasksVansh Kapoor, Aman Gupta, Hao Chen, Anurag Beniwal 等ICLR 2026 · 被引用 7 次
