SC2023Top-tier venue
Fine-grained Policy-driven I/O Sharing for Burst Buffers
Ed Karrels, Lei Huang, Yuhong Kan, Ishank Arora, Yinzhi Wang, Daniel S. Katz, William Gropp, Zhao Zhang
Abstract
A burst buffer is a common method to bridge the performance gap between the I/O needs of modern supercomputing applications and the performance of the shared file system on large-scale supercomputers. However, existing I/O sharing methods require resource isolation, offline profiling, or repeated execution that significantly limit the utilization and applicability of these systems. Here we present ThemisIO, a policy-driven I/O sharing framework for a remote-shared burst buffer: a dedicated group of I/O nodes, each with a local storage device. ThemisIO preserves high utilization by implementing opportunity fairness so that it can reallocate unused I/O resources to other applications. ThemisIO accurately and efficiently allocates I/O cycles among applications, purely based on real-time I/O behavior without requiring user-supplied information or offline-profiled application characteristics. ThemisIO supports a variety of fair sharing policies, such as user-fair, size-fair, as well as composite policies, e.g., group-then-user-fair. All these features are enabled by its statistical token design. ThemisIO can alter the execution order of incoming I/O requests based on assigned tokens to precisely balance I/O cycles between applications via time slicing, thereby enforcing processing isolation. Experiments using I/O benchmarks show that ThemisIO sustains 13.5--13.7% higher I/O throughput and 19.5--40.4% lower performance variation than existing algorithms. For real applications, ThemisIO significantly reduces the slowdown by 59.1--99.8% caused by I/O interference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca3f9376-6a68-4396-b6bf-599e02fcd6ceBuilds on1
Related papers
- R2B: high-efficiency and fair I/O scheduling for multi-tenant with differentiated demandsDiansen Sun, Yunpeng Chai, Chaoyang Liu, Weihao Sun et al.DAC 2022 · 2 citations
- Dynamic Chip Clustering and Task Allocation for Real-time FlashGyeongtaek Kim, Sungjin Lee, Hoon Sung ChwaDAC 2021 · 1 citation
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang et al.EuroSys 2024 · 13 citations
- HadaFS: A File System Bridging the Local and Shared Burst Buffer for Exascale SupercomputersXiaobin He, Bin Yang, Jie Gao, Wei Xiao et al.FAST 2023 · 24 citations
- Combining Buffered I/O and Direct I/O in Distributed File SystemsYingjin Qian, Marc-André Vef, Patrick Farrell, Andreas Dilger et al.FAST 2024 · 11 citations
