Tango: A Cross-layer Approach to Managing I/O Interference over Local Ephemeral Storage
Zhenbo Qiao, Qirui Tian, Zhenlu Qin, Jinzhen Wang, Qing Liu, Norbert Podhorszki, Scott Klasky, Hongjian Zhu
摘要
As simulation-based scientific discovery advances to exascale, a major question that the community is striving to answer is how to co-design data storage and complex physicsrich analytics in a way that the time to knowledge can be minimized for post-processing. A particular challenge is how to accommodate a broad spectrum of data analytics needsparticularly those that become clear only until very late during the post-processing, a scenario where existing methods, such as in situ processing, are unable or less effective in supporting data analytics. As HPC storage systems have become deeper and more complex with the recent addition of NVMe, die-stacked memory, and burst buffer, it requires fundamentally rethinking new paradigms and methods for data storage and analysis. This paper aims to address the issue of I/O interference for data analytics over local ephemeral storage, which is shared by multiple applications in a non-exclusive node usage scenario-often configured for small- to medium-sized clusters. At the core of this work is a coordinated cross-layer approach that reacts to storage interference from both storage and application layers. By decomposing and distributing analysis data across the storage hierarchy, data analytics can adapt to the interference by reducing or completely avoiding access to lower tiers whenever there is a high interference, while maintaining a prescribed error bound to limit the information loss. Meanwhile, proper actions are also taken at the storage layer to ensure sufficient bandwidth is allocated for retrieving an augmentation, which is based upon the cardinality and accuracy of the augmentation as well as the nature of an application. We evaluate three realworld data analytics, XGC, GenASiS, and CFD, on Chameleon, and quantitatively demonstrate that the I/O performance can be vastly improved, e.g., by 52% versus no adaptivity and 36% versus single-layer adaptivity, while maintaining acceptable outcomes of data analysis.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Taming I/O variation on QoS-less HPC storage: what can applications do?Zhenbo Qiao, Qing Liu, Norbert Podhorszki, Scott Klasky 等SC 2020 · 被引用 10 次
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min 等ASPLOS 2023 · 被引用 48 次
- MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive WorkloadsLuke Logan, Anthony Kougkas, Xian-He SunSC 2024 · 被引用 3 次
- Graphite: A NUMA-aware HPC System for Graph Analytics Based on a new MPI * X Parallelism ModelMohammad Hasanzadeh-Mofrad, Rami G. Melhem, Muhammad Yousuf Ahmad, Mohammad HammoudVLDB 2020 · 被引用 142 次
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie 等HPDC 2022 · 被引用 26 次
