T-TAMER: Provably Taming Trade-offs in ML Serving
Yuanyuan Yang, Ruimin Zhang, Jamie Morgenstern, Haifeng Xu
摘要
As machine learning models continue to grow in size and complexity, efficient serving faces increasingly broad trade-offs spanning accuracy, latency, resource usage, and other objectives. Multi-model serving further complicates these trade-offs; for example, in cascaded models, each early-exit decision balances latency reduction against potential accuracy loss. Despite the pervasiveness and importance of such trade-offs, current strategies remain largely heuristic and case-specific, limiting both their theoretical guarantees and general applicability.
We present a general framework, T-Tamer, which formalizes this setting as a multi-stage decision process, where the objective is to determine both when to exit and which model to consult. Our main result shows that recall (i.e., the ability to revisit earlier models) is both necessary and sufficient for achieving provable performance guarantees. In particular, we prove that strategies without recall cannot obtain any constant-factor approximation to the optimal trade-off, whereas recall-based strategies provably attain the optimal trade-off in polynomial time.
We validate our analysis through experiments on synthetic datasets and early-exit workloads for vision and NLP benchmarks. The results show that recall-based strategies consistently yield efficient accuracy–latency trade-offs. We hope this work provides a principled foundation for bridging heuristic practice with theoretical guarantees in the design of early-exit and cascaded models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraintWei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang 等ICML 2024 · 被引用 346 次
- Revenue Maximization for Query PricingShuchi Chawla, Shaleen Deep, Paraschos Koutris, Yifeng TengVLDB 2020 · 被引用 58 次
- Pandora's Box with Correlations: Learning and ApproximationShuchi Chawla, Evangelia Gergatsouli, Yifeng Teng, Christos Tzamos 等FOCS 2020 · 被引用 29 次
- Cost-aware Bayesian Optimization via the Pandora's Box Gittins IndexQian Xie, Raul Astudillo, Peter I. Frazier, Ziv Scully 等NeurIPS 2024 · 被引用 23 次
相关 Paper
- Cascadia: An Efficient Cascade Serving System for Large Language ModelsYouhe Jiang, Fangcheng Fu, Wanru Zhao, Stephan Rabanser 等ICLR 2026 · 被引用 7 次
- Improving DNN Inference Throughput Using Practical, Per-Input Compute AdaptationAnand Padmanabha Iyer, Mingyu Guan, Yinwei Dai, Rui Pan 等SOSP 2024 · 被引用 1 次
- Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML ServingYinwei Dai, Rui Pan, Anand P. Iyer, Kai Li 等SOSP 2024 · 被引用 4 次
- SATER: A Self-Aware and Token-Efficient Approach to Routing and CascadingYuanzhe Shen, Yide Liu, Zisu Huang, Ruicheng Yin 等EMNLP 2025
- A Unified Approach to Routing and Cascading for LLMsJasper Dekoninck, Maximilian Baader, Martin T. VechevICML 2025
