USENIX ATC2025顶会
On-Demand Container Partitioning for Distributed ML
Giovanni Bartolomeo, Navidreza Asadi, Wolfgang Kellerer, Jörg Ott, Nitinder Mohan
摘要
As machine learning (ML) models grow in complexity and scale, distributed deployment across multiple devices has become essential for ensuring performance and scalability. However, the dynamic nature of distributed ML, where models must be frequently retrained, partitioned, and updated, exposes severe limitations in the current de-facto containerbased model deployment. Specifically, the layered architecture of container filesystems is not well-suited for handling fine-grained model updates and partitioned ML deployments, leading to inefficient rebuilds and long delays. In this paper, we present 2DFS, a novel two-dimensional filesystem that enables independent updates, caching, and distribution of ML model components. We design and develop a complete ecosystem, including a builder, registry, and cache hierarchy, to streamline the build and deployment processes of ML models leveraging 2DFS. Our comprehensive evaluation of 14 real-world ML models demonstrates that 2DFS achieves up to 56x faster build times, 25x better caching efficiency, while providing on-demand image partitioning with negligible overhead. 2DFS is fully OCI-compliant and integrates seamlessly with existing infrastructures and container workflows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Unprecedented ScaleReza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li 等SC 2022 · 被引用 276 次
相关 Paper
- FalconFS: Distributed File System for Large-Scale Deep Learning PipelineJingwei Xu, Junbin Kang, Mingkai Dong, Mingyu Liu 等NSDI 2026 · 被引用 2 次
- Clairvoyant prefetching for distributed machine learning I/ONikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten HoeflerSC 2021 · 被引用 60 次
- Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With SenecaOmkar Desai, Ziyang Jiao, Shuyi Pei, Janki Bhimani 等FAST 2026 · 被引用 3 次
- FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud CommunicationJoe Oakley, Hakan FerhatosmanogluICDE 2024 · 被引用 6 次
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
