Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples
Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong
摘要
The rapid advancement of computer system components has necessitated a comprehensive profiling approach for both on-CPU and off-CPU events simultaneously. However, the conventional approach lacks profiling both on-and off-CPU events, so they fall short of accurately assessing the overhead of each bottleneck in modern applications.
In this paper, we propose a sampling-based profiling technique called blocked samples that is designed to capture all types of off-CPU events, such as I/O waiting, blocking synchronization, and waiting in CPU runqueue. Using the blocked samples technique, this paper proposes two profilers, bperf and BCOZ. Leveraging blocked samples, bperf profiles applications by providing symbol-level profile information when a thread is either on the CPU or off the CPU, awaiting scheduling or I/O requests. Using the information, BCOZ performs causality analysis of collected on-and off-CPU events to precisely identify performance bottlenecks and the potential impact of optimizations. The profiling capability of BCOZ is verified using real applications. From our profiling results followed by actual optimization, BCOZ identifies bottlenecks with off-CPU events precisely, and their optimization results are aligned with the predicted performance improvement by BCOZ's causality analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma 等FAST 2021 · 被引用 120 次
- SplinterDB: Closing the Bandwidth Gap for NVMe Key-Value StoresAlexander Conway, Abhishek Gupta, Vijay Chidambaram, Martin Farach-Colton 等USENIX ATC 2020 · 被引用 90 次
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 被引用 52 次
- ListDB: Union of Write-Ahead Logs and Persistent SkipLists for Incremental Checkpointing on Persistent MemoryWonbae Kim, Chanyeol Park, Dongui Kim, Hyeongjun Park 等OSDI 2022 · 被引用 47 次
- CruiseDB: An LSM-Tree Key-Value Store with Both Better Tail Throughput and Tail LatencyJunkai Liang, Yunpeng ChaiICDE 2021 · 被引用 17 次
相关 Paper
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh 等NSDI 2024 · 被引用 4 次
- MemPerf: Profiling Allocator-Induced Performance SlowdownsJin Zhou, Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang 等OOPSLA 2023
- Divining Profiler Accuracy: An Approach to Approximate Profiler Accuracy through Machine Code-Level SlowdownHumphrey Burchell, Stefan MarrOOPSLA 2025 · 被引用 3 次
- vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized CloudsWeiwei Jia, Jianchen Shan, Tsz On Li, Xiaowei Shang 等USENIX ATC 2020 · 被引用 16 次
- A Finer-Grained Blocking Analysis for Parallel Real-Time Tasks with Spin-LocksZe-Wei Chen, Hang Lei, Maolin Yang, Yong Liao 等DAC 2021 · 被引用 5 次
