EvoStore: Towards Scalable Storage of Evolving Learning Models
Robert Underwood, Meghana Madhyastha, Randal C. Burns, Bogdan Nicolae
Abstract
Deep Learning (DL) has seen rapid adoption in all domains. Since training DL models is expensive, both in terms of time and resources, application workflows that make use of DL increasingly need to operate with a large number of derived learning models, which are obtained through transfer learning and fine-tuning. At scale, thousands of such derived DL models are accessed concurrently by a large number of processes. In this context, an important question is how to design and develop specialized DL model repositories that remain scalable under concurrent access, while addressing key challenges: how to query the DL model architectures for specific patterns? How to load/store a subset of layers/tensors from a DL model? How to efficiently share unmodified layers/tensors between DL models derived from each other through transfer learning? How to maintain provenance and answer ancestry queries? State of art leaves a gap regarding these challenges. To fill this gap, we introduce EvoStore, a distributed DL model repository with scalable data and metadata support to store and access derived DL models efficiently. Large-scale experiments on hundreds of GPUs show significant benefits over state-of-art with respect to I/O and metadata performance, as well as storage space utilization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32278273-d336-40d7-85e2-e68364e55e51Builds on5
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
- Few-Shot Neural Architecture SearchYiyang Zhao, Linnan Wang, Yuandong Tian, Rodrigo Fonseca et al.ICML 2021 · 100 citations
- AgEBO-tabular: joint neural architecture and hyperparameter search with autotuned data-parallel training for tabular dataRomain Égelé, Prasanna Balaprakash, Isabelle Guyon, Venkatram Vishwanath et al.SC 2021 · 14 citations
- Check-N-Run: a Checkpointing System for Training Deep Learning Recommendation ModelsAssaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere et al.NSDI 2022
Related papers
- MGit: A Model Versioning and Management SystemWei Hao, Daniel Mendoza, Rafael Mendes, Deepak Narayanan et al.ICML 2024 · 1 citation
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and InferenceSai Wu, Ruichen Xia, Dingyu Yang, Rui Wang et al.SIGMOD 2026 · 1 citation
- A comprehensive study on challenges in deploying deep learning based softwareZhenpeng Chen, Yanbin Cao, Yuanqiang Liu, Haoyu Wang et al.FSE 2020 · 121 citations
- ModularEvo: Evolving Multi-Task Models via Neural Network Modularization and CompositionWenrui Long, Binhang Qi, Hailong Sun, Zongzhen Yang et al.ICSE 2026
- An Empirical Study of Pre-Trained Model Reuse in the Hugging Face Deep Learning Model RegistryWenxin Jiang, Nicholas Synovic, Matt Hyatt, Taylor R. Schorlemmer et al.ICSE 2023 · 62 citations
