USENIX ATC2023顶会
Tectonic-Shift: A Composite Storage Fabric for Large-Scale ML Training
Mark Zhao, Satadru Pan, Niket Agarwal, Zhaoduo Wen, David Xu, Anand Natarajan, Pavan Kumar, Shiva Shankar P., Ritesh Tijoriwala, Karan Asher, Hao Wu, Aarti Basant
摘要
Tectonic-Shift is the storage fabric for Meta's production machine learning (ML) training infrastructure. Industrial storage fabrics for ML need to meet both the intensive IO and highcapacity storage demands of training jobs. Our prior storage fabric, Tectonic, used hard disk drives (HDDs) to store training data. However, HDDs provide poor IO-per-watt performance. This inefficiency hindered the scalability of our storage fabric, and thus limited our ability to keep pace with rapidly growing training IO demands.
This paper describes our journey to build and deploy Tectonic-Shift, a composite storage fabric that efficiently serves the needs of our training infrastructure. We begin with a deep workload characterization that guided an extensive hardware and software design space exploration. We then present the principled design of Tectonic-Shift, which maximizes storage power efficiency by combining Shift, a flash storage tier, with Tectonic. Shift improves efficiency by absorbing reads using IO-efficient flash, reducing required HDD capacity. Shift maximizes IO absorption via novel applicationaware cache policies that infer future access patterns from training dataset specifications. Shift absorbs 1.51 -3.28× more IO than an LRU flash cache and reduces power demand in a petabyte-scale production Tectonic-Shift cluster by 29%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Baleen: ML Admission & Prefetching for Flash CachesDaniel Lin-Kit Wong, Hao Wu, Carson Molder, Sathya Gunasekar 等FAST 2024 · 被引用 26 次
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 被引用 13 次
- Cloudscape: A Study of Storage Services in Modern Cloud ArchitecturesSambhav Satija, Chenhao Ye, Ranjitha Kosgi, Aditya Jain 等FAST 2025 · 被引用 12 次
- PreSto: An In-Storage Data Preprocessing System for Training Recommendation ModelsYunjae Lee, Hyeseong Kim, Minsoo RhuISCA 2024 · 被引用 8 次
- MinatoLoader: Accelerating Machine Learning Training Through Efficient Data PreprocessingRahma Nouaji, Stella Bitchebe, Ricardo Macedo, Oana BalmauEuroSys 2026 · 被引用 3 次
它引用的顶会 Paper11
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 被引用 193 次
- The CacheLib Caching Engine: Design and Experiences at ScaleBenjamin Berg, Daniel S. Berger, Sara McAllister, Isaac Grosof 等OSDI 2020 · 被引用 145 次
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Facebook's Tectonic Filesystem: Efficiency from ExascaleSatadru Pan, Theano Stavrinos, Yunqiao Zhang, Atul Sikaria 等FAST 2021 · 被引用 110 次
- An Imitation Learning Approach for Cache ReplacementEvan Zheran Liu, Milad Hashemi, Kevin Swersky, Parthasarathy Ranganathan 等ICML 2020 · 被引用 108 次
相关 Paper
- Automating Distributed Tiered Storage Management in Cluster ComputingHerodotos Herodotou, Elena KakoulliVLDB 2020 · 被引用 30 次
- Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadJean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Bing Xie 等HPDC 2022 · 被引用 26 次
- Opus: Photonic Rail-Optimized Fabric in ML DatacentersEric Ding, Barry Lyu, Bhaskar Kataria, Rachee SinghSIGCOMM 2026
- Heimdall: Optimizing Storage I/O Admission with Extensive Machine Learning PipelineDaniar Heri Kurniawan, Rani Ayu Putri, Peiran Qin, Kahfi S. Zulkifli 等EuroSys 2025 · 被引用 3 次
- MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at HyperscaleArnab Choudhury, Yang Wang, Tuomas Pelkonen, Kutta Srinivasan 等OSDI 2024 · 被引用 39 次
