HadaFS: A File System Bridging the Local and Shared Burst Buffer for Exascale Supercomputers
Xiaobin He, Bin Yang, Jie Gao, Wei Xiao, Qi Chen, Shupeng Shi, Dexun Chen, Weiguo Liu, Wei Xue, Zuoning Chen
Abstract
Current supercomputers introduce SSDs to form a Burst Buffer (BB) layer to meet the HPC application's growing I/O requirements. BBs can be divided into two types by deployment location. One is the local BB, which is known for its scalability and performance. The other is the shared BB, which has the advantage of data sharing and deployment costs. How to unify the advantages of the local BB and the shared BB is a key issue in the HPC community.
We propose a novel BB file system named HadaFS that provides the advantages of local BB deployments to shared BB deployments. First, HadaFS offers a new Localized Triage Architecture (LTA) to solve the problem of ultra-scale expansion and data sharing. Then, HadaFS proposes a full-path indexing approach with three metadata synchronization strategies to solve the problem of complex metadata management of traditional file systems and mismatch with the application I/O behaviors. Moreover, HadaFS integrates a data management tool named Hadash, which supports efficient data query in the BB and accelerates data migration between the BB and traditional HPC storage. HadaFS has been deployed on the Sunway New-generation Supercomputer (SNS), serving hundreds of applications and supporting a maximum of 600,000-client scaling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85a4f130-d2e1-42d8-8be5-32af1952e9e7Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie et al.HPDC 2022 · 26 citations
- File System Semantics Requirements of HPC ApplicationsChen Wang, Kathryn Mohror, Marc SnirHPDC 2021 · 25 citations
Related papers
- Fine-grained Policy-driven I/O Sharing for Burst BuffersEd Karrels, Lei Huang, Yuhong Kan, Ishank Arora et al.SC 2023 · 5 citations
- DeltaFS: a scalable no-ground-truth filesystem for massively-parallel computingQing Zheng, Charles D. Cranor, Gregory R. Ganger, Garth A. Gibson et al.SC 2021 · 5 citations
- Concealing Compression-accelerated I/O for HPC Applications through In Situ Task SchedulingSian Jin, Sheng Di, Frédéric Vivien, Daoce Wang et al.EuroSys 2024 · 13 citations
- HAVS: Hardware-accelerated Shared-memory-based VPP Network StackShujun Zhuang, Jian Zhao, Jian Li, Ping Yu et al.INFOCOM 2021 · 1 citation
- BAASH: lightweight, efficient, and reliable blockchain-as-a-service for HPC systemsAbdullah Al-Mamun, Feng Yan, Dongfang ZhaoSC 2021 · 14 citations
