T-TAMER: Provably Taming Trade-offs in ML Serving
Yuanyuan Yang, Ruimin Zhang, Jamie Morgenstern, Haifeng Xu
Abstract
As machine learning models continue to grow in size and complexity, efficient serving faces increasingly broad trade-offs spanning accuracy, latency, resource usage, and other objectives. Multi-model serving further complicates these trade-offs; for example, in cascaded models, each early-exit decision balances latency reduction against potential accuracy loss. Despite the pervasiveness and importance of such trade-offs, current strategies remain largely heuristic and case-specific, limiting both their theoretical guarantees and general applicability.
We present a general framework, T-Tamer, which formalizes this setting as a multi-stage decision process, where the objective is to determine both when to exit and which model to consult. Our main result shows that recall (i.e., the ability to revisit earlier models) is both necessary and sufficient for achieving provable performance guarantees. In particular, we prove that strategies without recall cannot obtain any constant-factor approximation to the optimal trade-off, whereas recall-based strategies provably attain the optimal trade-off in polynomial time.
We validate our analysis through experiments on synthetic datasets and early-exit workloads for vision and NLP benchmarks. The results show that recall-based strategies consistently yield efficient accuracy–latency trade-offs. We hope this work provides a principled foundation for bridging heuristic practice with theoretical guarantees in the design of early-exit and cascaded models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4c45dab-d04f-4aaa-ad7a-d357ff49b76eBuilds on21
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang et al.ICML 2024 · 346 citations
- Revenue Maximization for Query PricingShuchi Chawla, Shaleen Deep, Paraschos Koutris, Yifeng TengVLDB 2020 · 58 citations
- Pandora's Box with Correlations: Learning and ApproximationShuchi Chawla, Evangelia Gergatsouli, Yifeng Teng, Christos Tzamos et al.FOCS 2020 · 29 citations
- Cost-aware Bayesian Optimization via the Pandora's Box Gittins IndexQian Xie, Raul Astudillo, Peter I. Frazier, Ziv Scully et al.NeurIPS 2024 · 23 citations
Related papers
- Cascadia: An Efficient Cascade Serving System for Large Language ModelsYouhe Jiang, Fangcheng Fu, Wanru Zhao, Stephan Rabanser et al.ICLR 2026 · 7 citations
- Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationAnand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan et al.SOSP 2024 · 1 citation
- Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML ServingYinwei Dai, Rui Pan, Anand P. Iyer, Kai Li et al.SOSP 2024 · 4 citations
- SATER: A Self-Aware and Token-Efficient Approach to Routing and CascadingYuanzhe Shen, Yide Liu, Zisu Huang, Ruicheng Yin et al.EMNLP 2025
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
