throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference Serving
Andreas Kosmas Kakolyris, Dimosthenis Masouros, Petros Vavaroutsos, Sotirios Xydis, Dimitrios Soudris
摘要
As Large Language Models (LLMs) gain traction, their reliance on power-hungry GPUs places ever-increasing energy demands, raising environmental and monetary concerns. Inference dominates LLM workloads, presenting a critical challenge for providers: minimizing energy costs under Service-Level Objectives (SLOs) that ensure optimal user experience. In this paper, we present throttLL’eM, a framework that reduces energy consumption while meeting SLOs through the use of instance and GPU frequency scaling. throttLL’eM features mechanisms that project future Key-Value (KV) cache usage and batch size. Leveraging a Machine-Learning (ML) model that receives these projections as inputs, throttLL’eM manages performance at the iteration level to satisfy SLOs with reduced frequencies and instance sizes. We show that the proposed ML model achieves scores greater than 0.97 and miss-predicts performance by less than 1 iteration per second on average. Experimental results on LLM inference traces show that throttLL’eM achieves up to lower energy consumption and an energy efficiency improvement of at least under SLOs, when compared to NVIDIA’s Triton server. throttLL’eM is publicly available at https://github.com/WilliamBlaskowicz/throttLL-eM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIMQingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang 等ISCA 2026 · 被引用 6 次
- LithOS: An Operating System for Efficient Machine Learning on GPUsPatrick H. Coppock, Brian Zhang, Eliot H. Solomon, Vasilis Kypriotis 等SOSP 2025 · 被引用 4 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
它引用的顶会 Paper29
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 被引用 14 次
- Energy Considerations of Large Language Model Inference and Efficiency OptimizationsJared Fernandez, Clara Na, Vashisth Tiwari, Yonatan Bisk 等ACL 2025
- TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceChenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao 等AAAI 2026 · 被引用 12 次
- TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud PlatformsJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Esha Choukse 等ASPLOS 2025 · 被引用 30 次
