TEA: Time-Proportional Event Analysis
Björn Gottschall, Lieven Eeckhout, Magnus Jahre
摘要
As computer architectures become increasingly complex and heterogeneous, it becomes progressively more difficult to write applications that make good use of hardware resources. Performance analysis tools are hence critically important as they are the only way through which developers can gain insight into the reasons why their application performs as it does. State-of-the-art performance analysis tools capture a plethora of performance events and are practically non-intrusive, but performance optimization is still extremely challenging. We believe that the fundamental reason is that current state-of-the-art tools in general cannot explain why executing the application's performance-critical instructions take time.
We hence propose Time-Proportional Event Analysis (TEA) which explains why the architecture spends time executing the application's performance-critical instructions by creating timeproportional Per-Instruction Cycle Stacks (PICS). PICS unify performance profiling and performance event analysis, and thereby (i) report the contribution of each static instruction to overall execution time, and (ii) break down per-instruction execution time across the (combinations of) performance events that a static instruction was subjected to across its dynamic executions. Creating timeproportional PICS requires tracking performance events across all in-flight instructions, but TEA only increases per-core power consumption by ∼3.2 mW (∼0.1%) because we carefully select events to balance insight and overhead. TEA leverages statistical sampling to keep performance overhead at 1.1% on average while incurring an average error of 2.1% compared to a non-sampling golden reference; a significant improvement upon the 55.6%, 55.5%, and 56.0% average error for AMD IBS, Arm SPE, and IBM RIS. We demonstrate that TEA's accuracy matters by using TEA to identify performance issues in the SPEC CPU2017 benchmarks lbm and nab that, once addressed, yield speedups of 1.28× and 2.45×, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li 等ASPLOS 2025 · 被引用 4 次
- Temporarily Unauthorized Stores: Write First, Ask for Permission LaterJuan M. Cebrian, Magnus Jahre, Alberto RosMICRO 2024 · 被引用 2 次
- Tintin: A Unified Hardware Performance Profiling Infrastructure to Uncover and Manage UncertaintyAo Li, Marion Sudvarg, Zihan Li, Sanjoy K. Baruah 等OSDI 2025 · 被引用 2 次
- Performance Predictability in Heterogeneous MemoryJinshu Liu, Hanchen Xu, Daniel S. Berger, Marcos K. Aguilera 等ASPLOS 2026 · 被引用 1 次
- Chips Need DIP: Time-Proportional Per-Instruction Cycle Stacks at DispatchSilvio Heverton Campelo de Santana, Joseph Rogers, Lieven Eeckhout, Magnus JahreASPLOS 2026 · 被引用 1 次
它引用的顶会 Paper3
- FirePerf: FPGA-Accelerated Full-System Hardware/Software Performance Profiling and Co-DesignSagar Karandikar, Albert J. Ou, Alon Amid, Howard Mao 等ASPLOS 2020 · 被引用 18 次
- TIP: Time-Proportional Instruction ProfilingBjörn Gottschall, Lieven Eeckhout, Magnus JahreMICRO 2021 · 被引用 16 次
- BayesPerf: minimizing performance monitoring errors using Bayesian statisticsSubho S. Banerjee, Saurabh Jha, Zbigniew Kalbarczyk, Ravishankar K. IyerASPLOS 2021 · 被引用 14 次
相关 Paper
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh 等NSDI 2024 · 被引用 4 次
- PerFlow: a domain specific framework for automatic performance analysis of parallel applicationsYuyang Jin, Haojie Wang, Runxin Zhong, Chen Zhang 等PPoPP 2022 · 被引用 10 次
- VESTA: Power Modeling with Language Runtime EventsJoseph Raskind, Timur Babakol, Khaled Mahmoud, Yu David LiuPLDI 2024 · 被引用 5 次
- DrCCTProf: a fine-grained call path profiler for ARM-based clustersQidong Zhao, Xu Liu, Milind ChabbiSC 2020 · 被引用 11 次
- Learning Generalizable Program and Architecture Representations for Performance ModelingLingda Li, Thomas Flynn, Adolfy HoisieSC 2024 · 被引用 5 次
