Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples
Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong
Abstract
The rapid advancement of computer system components has necessitated a comprehensive profiling approach for both on-CPU and off-CPU events simultaneously. However, the conventional approach lacks profiling both on-and off-CPU events, so they fall short of accurately assessing the overhead of each bottleneck in modern applications.
In this paper, we propose a sampling-based profiling technique called blocked samples that is designed to capture all types of off-CPU events, such as I/O waiting, blocking synchronization, and waiting in CPU runqueue. Using the blocked samples technique, this paper proposes two profilers, bperf and BCOZ. Leveraging blocked samples, bperf profiles applications by providing symbol-level profile information when a thread is either on the CPU or off the CPU, awaiting scheduling or I/O requests. Using the information, BCOZ performs causality analysis of collected on-and off-CPU events to precisely identify performance bottlenecks and the potential impact of optimizations. The profiling capability of BCOZ is verified using real applications. From our profiling results followed by actual optimization, BCOZ identifies bottlenecks with off-CPU events precisely, and their optimization results are aligned with the predicted performance improvement by BCOZ's causality analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bfc54f0-4673-4e3c-8b01-492ba7e1cc4eBuilds on7
- SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid StorageHao Chen, Chaoyi Ruan, Cheng Li, Xiaosong Ma et al.FAST 2021 · 120 citations
- SplinterDB: Closing the Bandwidth Gap for NVMe Key-Value StoresAlexander Conway, Abhishek Gupta, Vijay Chidambaram, Martin Farach-Colton et al.USENIX ATC 2020 · 90 citations
- ADOC: Automatically Harmonizing Dataflow Between Components in Log-Structured Key-Value Stores for Improved PerformanceJinghuan Yu, Sam H. Noh, Young-ri Choi, Chun Jason XueFAST 2023 · 52 citations
- ListDB: Union of Write-Ahead Logs and Persistent SkipLists for Incremental Checkpointing on Persistent MemoryWonbae Kim, Chanyeol Park, Dongui Kim, Hyeongjun Park et al.OSDI 2022 · 47 citations
- CruiseDB: An LSM-Tree Key-Value Store with Both Better Tail Throughput and Tail LatencyJunkai Liang, Yunpeng ChaiICDE 2021 · 17 citations
Related papers
- LDB: An Efficient Latency Profiling Tool for Multithreaded ApplicationsInho Cho, Seo Jin Park, Ahmed Saeed, Mohammad Alizadeh et al.NSDI 2024 · 4 citations
- MemPerf: Profiling Allocator-Induced Performance SlowdownsJin Zhou, Sam Silvestro, Steven (Jiaxun) Tang, Hanmei Yang et al.OOPSLA 2023
- Divining Profiler Accuracy: An Approach to Approximate Profiler Accuracy through Machine Code-Level SlowdownHumphrey Burchell, Stefan MarrOOPSLA 2025 · 3 citations
- vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized CloudsWeiwei Jia, Jianchen Shan, Tsz On Li, Xiaowei Shang et al.USENIX ATC 2020 · 16 citations
- A Finer-Grained Blocking Analysis for Parallel Real-Time Tasks with Spin-LocksZe-Wei Chen, Hang Lei, Maolin Yang, Yong Liao et al.DAC 2021 · 5 citations
