Principles and Methodologies for Serial Performance Optimization
Sujin Park, Mingyu Guan, Xiang Cheng, Taesoo Kim
Abstract
Throughout the history of computer science, optimizing existing systems to achieve higher performance has been a longstanding aspiration. While the primary emphasis of this endeavor lies in reducing latency and increasing throughput, these two are closely intertwined, and answering the how question has remained a challenge, often relying on intuition and experience.
This paper introduces a systematic approach to optimizing sequential tasks, which are fundamental for overall performance. We define three principles-task removal, replacement, and reordering-and distill them into eight actionable methodologies: batching, caching, precomputing, deferring, relaxation, contextualization, hardware specialization, and layering. Our review of OSDI and SOSP papers over the past decade shows that these techniques, when taken together, comprehensively account for the observed sequential optimization strategies.
To illustrate the framework's practical value, we present two case studies: one on file and storage systems, and another analyzing kernel synchronization to uncover missed optimization opportunities. Furthermore, we introduce SysGPT, a fine-tuned GPT model trained on curated literature analysis, which offer context-aware performance suggestions. SysGPT's outputs are more specific and feasible than GPT-4's, aligning with core strategies from recent research without direct exposure, demonstrating its utility as an optimization assistant.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b168234b-d55e-4270-8f4b-cf53a70133eeCited by top-tier papers1
Ask how each one uses itBuilds on37
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- Protean: VM Allocation Service at ScaleOri Hadary, Luke Marshall, Ishai Menache, Abhisek Pan et al.OSDI 2020 · 189 citations
- XRP: In-Kernel Storage Functions with eBPFYuhong Zhong, Haoyu Li, Yu Jian Wu, Ioannis Zarkadas et al.OSDI 2022 · 100 citations
- HeMem: Scalable Tiered Memory Management for Big Data Applications and Real NVMAmanda Raybuck, Tim Stamler, Wei Zhang, Mattan Erez et al.SOSP 2021 · 93 citations
- VBASE: Unifying Online Vector Similarity Search and Relational Queries via Relaxed MonotonicityQianxi Zhang, Shuotao Xu, Qi Chen, Guoxin Sui et al.OSDI 2023 · 75 citations
Related papers
- HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End OptimizationSize Li, Zhiqing Tang, Hongrui Liang, Jianxiong Guo et al.ICML 2026
- vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized CloudsWeiwei Jia, Jianchen Shan, Tsz On Li, Xiaowei Shang et al.USENIX ATC 2020 · 16 citations
- The Art of Latency Hiding in Modern Database EnginesKaisong Huang, Tianzheng Wang, Qingqing Zhou, Qingzhong MengVLDB 2024 · 23 citations
- Scalable Far Memory: Balancing Faults and EvictionsYueyang Pan, Yash Lala, Musa Unal, Yujie Ren et al.SOSP 2025
- Rapid Data Ingestion through DB-OS Co-designKyungmin Lim, Minseok Yoon, Kihwan Kim, Alan David Fekete et al.SIGMOD 2025 · 1 citation
