Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy Efficiency
Zibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang, Yanlin Liu, Zhiheng Hu, Jingyi Zhang, Xiaoxin Xu, Jian He, Xiaoliang Wang, Wanchun Dou, Guihai Chen, Chen Tian
摘要
Recent advancements in deep learning have significantly increased AI processors' energy consumption, which is becoming a critical factor limiting AI development. Dynamic Voltage and Frequency Scaling (DVFS) stands as a key method in power optimization. However, due to the latency of DVFS control in AI processors, previous works typically apply DVFS control at the granularity of a program's entire duration or sub-phases, rather than at the level of AI operators.
The advent of millisecond-level DVFS capabilities on the latest Ascend NPU platforms enables us to set frequency
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
- Beyond Utilization: Energy-Conscious GPU Sharing for Inference ServingPrasoon Sinha, Dimitrios Liakopoulos, Nathan Lemma, Neeraja J. YadwadkarSOSP 2026
- Inference in the Shadows: Taming Memory Bandwidth Contention in Mobile LLM Inference with SerenoTong Xin, Xinrui Shi, Mingkai Dong, Zeyu MiOSDI 2026
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUsMarco Kurzynski, Shaizeen Aga, Di WuISCA 2026
它引用的顶会 Paper7
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan 等MICRO 2021 · 被引用 134 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- ExeGPT: Constraint-Aware Resource Scheduling for LLM InferenceHyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim 等ASPLOS 2024 · 被引用 41 次
- Improving GPU Energy Efficiency through an Application-transparent Frequency Scaling Policy with Performance AssuranceYijia Zhang, Qiang Wang, Zhe Lin, Pengxiang Xu 等EuroSys 2024 · 被引用 18 次
相关 Paper
- Predict; Don't React for Enabling Efficient Fine-Grain DVFS in GPUsSrikant Bharadwaj, Shomit Das, Kaushik Mazumdar, Bradford M. Beckmann 等ASPLOS 2023 · 被引用 15 次
- PowerWeave: Unlocking Energy-Efficient ML on GPUs with OS-Level Spatial Power ManagementVasilis Kypriotis, Eric Dubberstein, Patrick H. Coppock, Eliot H. Solomon 等ISCA 2026 · 被引用 1 次
- PowerLens: An Adaptive DVFS Framework for Optimizing Energy Efficiency in Deep Neural NetworksJiawei Geng, Zongwei Zhu, Weihong Liu, Xuehai Zhou 等DAC 2024 · 被引用 10 次
- Know Your Enemy To Save Cloud Energy: Energy-Performance Characterization of Machine Learning ServingJunyeol Yu, Jongseok Kim, Euiseong SeoHPCA 2023 · 被引用 14 次
- Squeezing Operator Performance Potential for the Ascend ArchitectureYuhang Zhou, Zhibin Wang, Guyue Liu, Shipeng Li 等ASPLOS 2025 · 被引用 3 次
