Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at Dispatch
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
摘要
Despite continuous accelerator improvements, high single-thread performance remains crucial due to Amdahl's Law. With technology scaling slowing, developers must produce code that is as performant as possible, which often involves instruction-level performance analysis. Such analysis can be visualized as Per-Instruction Cycle Stacks (PICS), which, when created at the commit stage of the processor, report each static instruction's contribution to overall execution time and capture the performance events it was subject to. Sadly, PICS at commit (PICSC) often cannot explain performance because they solely capture how effectively instructions egress from the processor core's out-of-order execution window. In contrast, developers must typically also gain insight into how efficiently instructions ingress into the execution window to fully understand application performance. PICSC must hence be complemented by PICS at dispatch (PICSD) because execution window ingress issues result in specific instructions exhibiting high dispatch latencies. We therefore propose dispatch-time profiling (DIP), which combines time-proportional at dispatch attribution policies with statistical sampling to accurately report each static instruction's contribution to dispatch time as well as list the reason(s) for delayed ingress. DIP is simple to implement and incurs low overhead, i.e., storage and execution time and overheads of 49 bytes and 1%, respectively, while delivering high accuracy (average profile error of 5.2%). This is a significant improvement over the 26.9% error of state-of-the-art dispatch-tagging as implemented in AMD IBS, Arm SPE, and IBM RIS. We demonstrate that needing PICSD and PICSC is the common case by showing that 18 out of our 22 SPEC2017 benchmarks simultaneously exhibit both ingress and egress issues. Additionally, we leverage DIP's PICSD to optimize the fotonik3d benchmark, improving its performance by 8.5%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Leaky Frontends: Security Vulnerabilities in Processor FrontendsShuwen Deng, Bowen Huang, Jakub SzeferHPCA 2022 · 被引用 27 次
- TIP: Time-Proportional Instruction ProfilingBjörn Gottschall, Lieven Eeckhout, Magnus JahreMICRO 2021 · 被引用 16 次
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 被引用 14 次
- Exploring Instruction Fusion Opportunities in General Purpose ProcessorsSawan Singh, Arthur Perais, Alexandra Jimborean, Alberto RosMICRO 2022 · 被引用 11 次
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 被引用 11 次
相关 Paper
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 被引用 5 次
- UDP: Utility-Driven Fetch Directed Instruction PrefetchingSurim Oh, Mingsheng Xu, Tanvir Ahmed Khan, Baris Kasikci 等ISCA 2024 · 被引用 12 次
- Understanding Accelerator Compilers via Performance ProfilingAyaka Yorihiro, Griffin Berlstein, Pedro Pontes García, Kevin Laeufer 等OOPSLA 2026
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh 等NSDI 2024 · 被引用 4 次
- VegaProf: Profiling Vega VisualizationsJunran Yang, Alex Bäuerle, Dominik Moritz, Çagatay DemiralpUIST 2023 · 被引用 5 次
