MGit: A Model Versioning and Management System
Wei Hao, Daniel Mendoza, Rafael Mendes, Deepak Narayanan, Amar Phanishayee, Asaf Cidon, Junfeng Yang
Abstract
Models derived from other models are extremely common in machine learning (ML) today. For example, transfer learning is used to create task-specific models from "pre-trained" models through finetuning. This has led to an ecosystem where models are related to each other, sharing structure and often even parameter values. However, it is hard to manage these model derivatives: the storage overhead of storing all derived models quickly becomes onerous, prompting users to get rid of intermediate models that might be useful for further analysis. Additionally, undesired behaviors in models are hard to track down (e.g., is a bug inherited from an upstream model?). In this paper, we propose a model versioning and management system called MGit that makes it easier to store, test, update, and collaborate on model derivatives. MGit introduces a lineage graph that records provenance and versioning information between models, optimizations to efficiently store model parameters, as well as abstractions over this lineage graph that facilitate relevant testing, updating and collaboration functionality. MGit is able to reduce the lineage graph's storage footprint by up to 7× and automatically update downstream models in response to updates to upstream models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21ee9a0b-cea2-4c13-b16a-4f665b6a123cCited by top-tier papers1
Ask how each one uses itBuilds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Adaptive Testing and Debugging of NLP ModelsMarco Túlio Ribeiro, Scott M. LundbergACL 2022 · 99 citations
- Evaluating the Robustness of Neural Language Models to Input PerturbationsMilad Moradi, Matthias SamwaldEMNLP 2021 · 64 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
Related papers
- Git-Theta: A Git Extension for Collaborative Development of Machine Learning ModelsNikhil Kandpal, Brian Lester, Mohammed Muqeeth, Anisha Mascarenhas et al.ICML 2023 · 15 citations
- EvoStore: Towards Scalable Storage of Evolving Learning ModelsRobert Underwood, Meghana Madhyastha, Randal C. Burns, Bogdan NicolaeHPDC 2024 · 3 citations
- MLCask: Efficient Management of Component Evolution in Collaborative Data Analytics PipelinesZhaojing Luo, Sai Ho Yeung, Meihui Zhang, Kaiping Zheng et al.ICDE 2021 · 31 citations
- Neural LineageRunpeng Yu, Xinchao WangCVPR 2024
- Model Selection with Model Zoo via Graph LearningZiyu Li, Hilco van der Wilk, Danning Zhan, Megha Khosla et al.ICDE 2024 · 6 citations
