Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at Dispatch
Silvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus Jahre
Abstract
Despite continuous accelerator improvements, high single-thread performance remains crucial due to Amdahl's Law. With technology scaling slowing, developers must produce code that is as performant as possible, which often involves instruction-level performance analysis. Such analysis can be visualized as Per-Instruction Cycle Stacks (PICS), which, when created at the commit stage of the processor, report each static instruction's contribution to overall execution time and capture the performance events it was subject to. Sadly, PICS at commit (PICSC) often cannot explain performance because they solely capture how effectively instructions egress from the processor core's out-of-order execution window. In contrast, developers must typically also gain insight into how efficiently instructions ingress into the execution window to fully understand application performance. PICSC must hence be complemented by PICS at dispatch (PICSD) because execution window ingress issues result in specific instructions exhibiting high dispatch latencies. We therefore propose dispatch-time profiling (DIP), which combines time-proportional at dispatch attribution policies with statistical sampling to accurately report each static instruction's contribution to dispatch time as well as list the reason(s) for delayed ingress. DIP is simple to implement and incurs low overhead, i.e., storage and execution time and overheads of 49 bytes and 1%, respectively, while delivering high accuracy (average profile error of 5.2%). This is a significant improvement over the 26.9% error of state-of-the-art dispatch-tagging as implemented in AMD IBS, Arm SPE, and IBM RIS. We demonstrate that needing PICSD and PICSC is the common case by showing that 18 out of our 22 SPEC2017 benchmarks simultaneously exhibit both ingress and egress issues. Additionally, we leverage DIP's PICSD to optimize the fotonik3d benchmark, improving its performance by 8.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c78f2d85-293e-481c-873f-7f45961bb3c8Builds on8
- Leaky Frontends: Security Vulnerabilities in Processor FrontendsShuwen Deng, Bowen Huang, Jakub SzeferHPCA 2022 · 27 citations
- TIP: Time-Proportional Instruction ProfilingBjörn Gottschall, Lieven Eeckhout, Magnus JahreMICRO 2021 · 16 citations
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 14 citations
- Exploring Instruction Fusion Opportunities in General Purpose ProcessorsSawan Singh, Arthur Perais, Alexandra Jimborean, Alberto RosMICRO 2022 · 11 citations
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 11 citations
Related papers
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 5 citations
- UDP: Utility-Driven Fetch Directed Instruction PrefetchingSurim Oh, Mingsheng Xu, Tanvir Ahmed Khan, Baris Kasikci et al.ISCA 2024 · 12 citations
- Understanding Accelerator Compilers via Performance ProfilingAyaka Yorihiro, Griffin Berlstein, Pedro Pontes García, Kevin Laeufer et al.OOPSLA 2026
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh et al.NSDI 2024 · 4 citations
- VegaProf: Profiling Vega VisualizationsJunran Yang, Alex Bäuerle, Dominik Moritz, Çagatay DemiralpUIST 2023 · 5 citations
