TIP: Time-Proportional Instruction Profiling
Björn Gottschall, Lieven Eeckhout, Magnus Jahre
Abstract
A fundamental part of developing software is to understand what the application spends time on. This is typically determined using a performance profiler which essentially captures how execution time is distributed across the instructions of a program. At the same time, the highly parallel execution model of modern highperformance processors means that it is difficult to reliably attribute time to instructions -resulting in performance analysis being unnecessarily challenging.
In this work, we first propose the Oracle profiler which is a golden reference for performance profilers. Oracle is golden because (i) it accounts every clock cycle and every dynamic instruction, and (ii) it is time-proportional, i.e., it attributes a clock cycle to the instruction(s) that the processor exposes the latency of. We use Oracle to, for the first time, quantify the error of softwarelevel profiling, the dispatch-tagging heuristic used in AMD IBS and Arm SPE, the Last-Committing Instruction (LCI) heuristic used in external monitors, and the Next-Committing Instruction (NCI) heuristic used in Intel PEBS, resulting in average instruction-level profile errors of 61.8%, 53.1%, 55.4%, and 9.3%, respectively. The reason for these errors is that all existing profilers have cases in which they systematically attribute execution time to instructions that are not the root cause of performance loss. To overcome this issue, we propose Time-Proportional Instruction Profiling (TIP) which combines Oracle's time attribution policies with statistical sampling to enable practical implementation. We implement TIP within the Berkeley Out-of-Order Machine (BOOM) and find that TIP is highly accurate. More specifically, TIP's instruction-level profile error is only 1.6% on average (maximally 5.0%) versus 9.3% on average (maximally 21.0%) for state-of-the-art NCI. TIP's improved accuracy matters in practice, as we exemplify by using TIP to identify a performance problem in the SPEC CPU2017 benchmark Imagick that, once addressed, improves performance by 1.93×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8680b6ca-f363-46e2-8929-45335fffb031Cited by top-tier papers6
- TEA: Time-Proportional Event AnalysisBjörn Gottschall, Lieven Eeckhout, Magnus JahreISCA 2023 · 11 citations
- FireAxe: Partitioned FPGA-Accelerated Simulation of Large-Scale RTL DesignsJoonho Whangbo, Edwin Lim, Chengyi Lux Zhang, Kevin Anderson et al.ISCA 2024 · 6 citations
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li et al.ASPLOS 2025 · 4 citations
- Divining Profiler Accuracy: An Approach to Approximate Profiler Accuracy through Machine Code-Level SlowdownHumphrey Burchell, Stefan MarrOOPSLA 2025 · 3 citations
- Tintin: A Unified Hardware Performance Profiling Infrastructure to Uncover and Manage UncertaintyAo Li, Marion Sudvarg, Zihan Li, Sanjoy K. Baruah et al.OSDI 2025 · 2 citations
Builds on2
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao et al.ASPLOS 2020 · 18 citations
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 14 citations
Related papers
- Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at DispatchSilvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus JahreASPLOS 2026 · 1 citation
- Profile inference revisitedWenlei He, Julián Mestre, Sergey Pupyrev, Lei Wang et al.POPL 2022 · 16 citations
- DrCCTProf: a fine-grained call path profiler for ARM-based clustersQidong Zhao, Xu Liu, Milind ChabbiSC 2020 · 11 citations
- APT-GET: profile-guided timely software prefetchingSaba Jamilan, Tanvir Ahmed Khan, Grant Ayers, Baris Kasikci et al.EuroSys 2022 · 25 citations
- Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML FusionArash Nasr-Esfahany, Mohammad Alizadeh, Victor Lee, Hanna Alam et al.ISCA 2025 · 1 citation
