TIP: Time-Proportional Instruction Profiling
Björn Gottschall, Lieven Eeckhout, Magnus Jahre
摘要
A fundamental part of developing software is to understand what the application spends time on. This is typically determined using a performance profiler which essentially captures how execution time is distributed across the instructions of a program. At the same time, the highly parallel execution model of modern highperformance processors means that it is difficult to reliably attribute time to instructions -resulting in performance analysis being unnecessarily challenging.
In this work, we first propose the Oracle profiler which is a golden reference for performance profilers. Oracle is golden because (i) it accounts every clock cycle and every dynamic instruction, and (ii) it is time-proportional, i.e., it attributes a clock cycle to the instruction(s) that the processor exposes the latency of. We use Oracle to, for the first time, quantify the error of softwarelevel profiling, the dispatch-tagging heuristic used in AMD IBS and Arm SPE, the Last-Committing Instruction (LCI) heuristic used in external monitors, and the Next-Committing Instruction (NCI) heuristic used in Intel PEBS, resulting in average instruction-level profile errors of 61.8%, 53.1%, 55.4%, and 9.3%, respectively. The reason for these errors is that all existing profilers have cases in which they systematically attribute execution time to instructions that are not the root cause of performance loss. To overcome this issue, we propose Time-Proportional Instruction Profiling (TIP) which combines Oracle's time attribution policies with statistical sampling to enable practical implementation. We implement TIP within the Berkeley Out-of-Order Machine (BOOM) and find that TIP is highly accurate. More specifically, TIP's instruction-level profile error is only 1.6% on average (maximally 5.0%) versus 9.3% on average (maximally 21.0%) for state-of-the-art NCI. TIP's improved accuracy matters in practice, as we exemplify by using TIP to identify a performance problem in the SPEC CPU2017 benchmark Imagick that, once addressed, improves performance by 1.93×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 被引用 11 次
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson 等ISCA 2024 · 被引用 6 次
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li 等ASPLOS 2025 · 被引用 4 次
- Divining Profiler Accuracy: An Approach to Approximate Profiler Accuracy through Machine Code-Level SlowdownHumphrey Burchell, Stefan MarrOOPSLA 2025 · 被引用 3 次
- Tintin: A Unified Hardware Performance Profiling Infrastructure to Uncover and Manage UncertaintyAo Li, Marion Sudvarg, Zihan Li, Sanjoy K. Baruah 等OSDI 2025 · 被引用 2 次
它引用的顶会 Paper2
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao 等ASPLOS 2020 · 被引用 18 次
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 被引用 14 次
相关 Paper
- Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at DispatchSilvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus JahreASPLOS 2026 · 被引用 1 次
- Profile inference revisitedWenlei He, Julián Mestre, Sergey Pupyrev, Lei Wang 等POPL 2022 · 被引用 16 次
- DrCCTProf: a fine-grained call path profiler for ARM-based clustersQidong Zhao, Xu Liu, Milind ChabbiSC 2020 · 被引用 11 次
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci 等EuroSys 2022 · 被引用 25 次
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam 等ISCA 2025 · 被引用 1 次
