USENIX ATC2021顶会
Fair Scheduling for AVX2 and AVX-512 Workloads
Mathias Gottschlag, Philipp Machauer, Yussuf Khalil, Frank Bellosa
摘要
CPU schedulers such as the Linux Completely Fair Scheduler try to allocate equal shares of the CPU performance to tasks of equal priority by allocating equal CPU time as a technique to improve quality of service for individual tasks. Recently, CPUs have, however, become power-limited to the point where different subsets of the instruction set allow for different operating frequencies depending on the complexity of the instructions. In particular, Intel CPUs with support for AVX2 and AVX-512 instructions often reduce their frequency when these 256-bit and 512-bit SIMD instructions are used in order to prevent excessive power consumption. This frequency reduction often impacts other less power-intensive processes, in which case equal allocation of CPU time results in unequal performance and a substantial lack of performance isolation.
We describe a modification to existing schedulers to restore fairness for workloads involving tasks which execute complex power-intensive instructions. In particular, we present a technique to identify AVX2/AVX-512 tasks responsible for frequency reduction, and we modify CPU time accounting to increase the priority of other tasks slowed down by these AVX2/AVX-512 tasks. Whereas previously non-AVX applications running in parallel to AVX-512 applications were slowed down by 24.9% on average, our prototype reduces the performance difference between non-AVX tasks and AVX-512 tasks in such scenarios to 5.4% on average, with a similar improvement for workloads involving AVX2 applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 被引用 23 次
- OS scheduling with nest: keeping tasks close together on warm coresJulia Lawall, Himadri Chhaya-Shailesh, Jean-Pierre Lozi, Baptiste Lepers 等EuroSys 2022 · 被引用 9 次
- Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionEdward Guo, Weiwei Jia, Xiaoning Ding, Jianchen ShanEuroSys 2025 · 被引用 3 次
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo 等HPCA 2026 · 被引用 2 次
相关 Paper
- Fewer Cores, More Hertz: Leveraging High-Frequency Cores in the OS Scheduler for Improved Application PerformanceRedha Gouicem, Damien Carver, Jean-Pierre Lozi, Julien Sopena 等USENIX ATC 2020 · 被引用 12 次
- Improving Predication Efficiency through Compaction/Restoration of SIMD InstructionsAdrián Barredo, Juan M. Cebrian, Miquel Moretó, Marc Casas 等HPCA 2020 · 被引用 8 次
- BlueFace: Integrating an Accelerator into the Core's Pipeline through Algorithm-Interface Co-Design for Real-Time SoCsZhe Jiang, Nathan Fisher, Nan Guan, Zheng DongDAC 2023 · 被引用 1 次
- MAFin: Maximizing Accuracy in FinFET based Approximated Real-Time ComputingShounak Chakraborty, Sangeet Saha, Magnus Själander, Klaus D. McDonald-MaierDAC 2024
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
