Modyn: Data-Centric Machine Learning Pipeline Orchestration
Maximilian Böther, Ties Robroek, Viktor Gsteiger, Robin Holzinger, Xianzhe Ma, Pinar Tözün, Ana Klimovic
Abstract
In real-world machine learning (ML) pipelines, datasets are continuously growing. Models must incorporate this new training data to improve generalization and adapt to potential distribution shifts. The cost of model retraining is proportional to how frequently the model is retrained and how much data it is trained on, which makes the naive approach of retraining from scratch each time impractical.
We present Modyn, a data-centric end-to-end machine learning platform. Modyn's ML pipeline abstraction enables users to declaratively describe policies for continuously training a model on a growing dataset. Modyn pipelines allow users to apply data selection policies (to reduce the number of data points) and triggering policies (to reduce the number of trainings). Modyn executes and orchestrates these continuous ML training pipelines. The system is open-source and comes with an ecosystem of benchmark datasets, models, and tooling. We formally discuss how to measure the performance of ML pipelines by introducing the concept of composite models, enabling fair comparison of pipelines with different data selection and triggering policies.
We empirically analyze how various data selection and triggering policies impact model accuracy, and also show that Modyn enables high throughput training with sample-level data selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb237b9b-d6da-441a-8b69-26ee26dad65dCited by top-tier papers4
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur et al.SIGMOD 2026 · 5 citations
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal et al.VLDB 2026
- RoCE BALBOA: Service-Enhanced RDMA Offload Engine for Data Center SmartNICsMaximilian Jakob Heer, Benjamin Ramhorst, Yu Zhu, Luhao Liu et al.OSDI 2026
- NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling: [Experiments & Analysis]Zhanhao Zhao, Haotian Gao, Naili Xing, Lingze Zeng et al.SIGMOD 2026
Builds on32
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- FILA: Online Auditing of Machine Learning Model Accuracy under Finite Labelling BudgetNaiqing Guan, Nick KoudasSIGMOD 2022 · 1 citation
- cedar: Optimized and Unified Machine Learning Input Data PipelinesMark Zhao, Emanuel Adamiak, Christos KozyrakisVLDB 2025 · 13 citations
- Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning PipelinesStefan Grafberger, Paul Groth, Sebastian SchelterSIGMOD 2023 · 18 citations
- Data Agent: Learning to Select Data via End-to-End Dynamic OptimizationSuorong Yang, Fangjian Su, Hai Gan, Ziqi Ye et al.ICML 2026
- Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning MultiverseJan Simson, Fiona Draxler, Samuel Mehr, Christoph KernCHI 2025 · 3 citations
