TEA: Time-Proportional Event Analysis
Björn Gottschall, Lieven Eeckhout, Magnus Jahre
Abstract
As computer architectures become increasingly complex and heterogeneous, it becomes progressively more difficult to write applications that make good use of hardware resources. Performance analysis tools are hence critically important as they are the only way through which developers can gain insight into the reasons why their application performs as it does. State-of-the-art performance analysis tools capture a plethora of performance events and are practically non-intrusive, but performance optimization is still extremely challenging. We believe that the fundamental reason is that current state-of-the-art tools in general cannot explain why executing the application's performance-critical instructions take time.
We hence propose Time-Proportional Event Analysis (TEA) which explains why the architecture spends time executing the application's performance-critical instructions by creating timeproportional Per-Instruction Cycle Stacks (PICS). PICS unify performance profiling and performance event analysis, and thereby (i) report the contribution of each static instruction to overall execution time, and (ii) break down per-instruction execution time across the (combinations of) performance events that a static instruction was subjected to across its dynamic executions. Creating timeproportional PICS requires tracking performance events across all in-flight instructions, but TEA only increases per-core power consumption by ∼3.2 mW (∼0.1%) because we carefully select events to balance insight and overhead. TEA leverages statistical sampling to keep performance overhead at 1.1% on average while incurring an average error of 2.1% compared to a non-sampling golden reference; a significant improvement upon the 55.6%, 55.5%, and 56.0% average error for AMD IBS, Arm SPE, and IBM RIS. We demonstrate that TEA's accuracy matters by using TEA to identify performance issues in the SPEC CPU2017 benchmarks lbm and nab that, once addressed, yield speedups of 1.28× and 2.45×, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73b96edd-68d4-43e3-afdd-c4c49c229401Cited by top-tier papers5
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li et al.ASPLOS 2025 · 4 citations
- Temporarily Unauthorized Stores: Write First, Ask for Permission LaterJuan M. Cebrian, Magnus Jahre, Alberto RosMICRO 2024 · 2 citations
- Tintin: A Unified Hardware Performance Profiling Infrastructure to Uncover and Manage UncertaintyAo Li, Marion Sudvarg, Zihan Li, Sanjoy K. Baruah et al.OSDI 2025 · 2 citations
- Performance Predictability in Heterogeneous MemoryJinshu Liu, Hanchen Xu, Daniel S. Berger, Marcos K. Aguilera et al.ASPLOS 2026 · 1 citation
- Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at DispatchSilvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus JahreASPLOS 2026 · 1 citation
Builds on3
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao et al.ASPLOS 2020 · 18 citations
- TIP: Time-Proportional Instruction ProfilingBjörn Gottschall, Lieven Eeckhout, Magnus JahreMICRO 2021 · 16 citations
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 14 citations
Related papers
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh et al.NSDI 2024 · 4 citations
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang et al.PPoPP 2022 · 10 citations
- VESTA: Power Modeling with Language Runtime EventsJoseph Raskind, Timur Babakol, Khaled Mahmoud, Yu David LiuPLDI 2024 · 5 citations
- DrCCTProf: a fine-grained call path profiler for ARM-based clustersQidong Zhao, Xu Liu, Milind ChabbiSC 2020 · 11 citations
- Learning Generalizable Program and Architecture Representations for Performance ModelingLingda Li, Thomas Flynn, Adolfy HoisieSC 2024 · 5 citations
