Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches
Yuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak, Ekta Khanna, Orhan Kislal, Domino Valdano, Arun Kumar
摘要
Deep learning (DL) is growing in popularity for many data analytics applications, including among enterprises. Large business-critical datasets in such settings typically reside in RDBMSs or other data systems. The DB community has long aimed to bring machine learning (ML) to DBMS-resident data. Given past lessons from in-DBMS ML and recent advances in scalable DL systems, DBMS and cloud vendors are increasingly interested in adding more DL support for DB-resident data. Recently, a new parallel DL model selection execution approach called Model Hopper Parallelism (MOP) was proposed. In this paper, we characterize the particular suitability of MOP for DL on data systems, but to bring MOP-based DL to DBresident data, we show that there is no single "best" approach, and an interesting tradeoff space of approaches exists. We explain four canonical approaches and build prototypes upon Greenplum Database, compare them analytically on multiple criteria (e.g., runtime efficiency and ease of governance) and compare them empirically with large-scale DL workloads. Our experiments and analyses show that it is non-trivial to meet all practical desiderata well and there is a Pareto frontier; for instance, some approaches are 3x-6x faster but fare worse on governance and portability. Our results and insights can help DBMS and cloud vendors design better DL support for DB users. All of our source code, data, and other artifacts are available at https://github.com/makemebitter/cerebro-ds .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- End-to-end Optimization of Machine Learning Prediction QueriesKwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen 等SIGMOD 2022 · 被引用 50 次
- User-Defined Operators: Efficiently Integrating Custom Algorithms into Modern DatabasesMoritz Sichert, Thomas NeumannVLDB 2022 · 被引用 23 次
- Scalable Graph Convolutional Network Training on Distributed-Memory SystemsGunduz Vehbi Demirci, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2023 · 被引用 18 次
- Database Native Model Selection: Harnessing Deep Neural Networks in Database SystemsNaili Xing, Shaofeng Cai, Gang Chen, Zhaojing Luo 等VLDB 2024 · 被引用 14 次
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 被引用 13 次
它引用的顶会 Paper5
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 被引用 61 次
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm 等SIGMOD 2020 · 被引用 27 次
- Optimizing Machine Learning Workloads in Collaborative EnvironmentsBehrouz Derakhshan, Alireza Rezaei Mahdiraji, Ziawasch Abedjan, Tilmann Rabl 等SIGMOD 2020 · 被引用 22 次
- Dynamic Parameter Allocation in Parameter ServersAlexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker MarklVLDB 2020 · 被引用 18 次
- Vista: Optimized System for Declarative Feature Transfer from Deep CNNs at ScaleSupun Nakandala, Arun KumarSIGMOD 2020 · 被引用 12 次
相关 Paper
- MorphingDB: A Task-Centric AI-Native DBMS for Model Management and InferenceSai Wu, Ruichen Xia, Dingyu Yang, Rui Wang 等SIGMOD 2026 · 被引用 1 次
- DistVec: Efficient Distributed Machine Learning in Parallel Database SystemsXinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu 等ICDE 2026
- Aero: Adaptive Query Processing of ML QueriesGaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj 等SIGMOD 2025 · 被引用 2 次
- Powering In-Database Dynamic Model Slicing for Structured Data AnalyticsLingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen 等VLDB 2024 · 被引用 7 次
- HYPPO: Using Equivalences to Optimize Pipelines in Exploratory Machine LearningAntonios Kontaxakis, Dimitris Sacharidis, Alkis Simitsis, Alberto Abelló 等ICDE 2024 · 被引用 1 次
