Distributed Numerical and Machine Learning Computations via Two-Phase Execution of Aggregated Join Trees
Dimitrije Jankov, Binhang Yuan, Shangyu Luo, Chris Jermaine
Abstract
When numerical and machine learning (ML) computations are expressed relationally, classical query execution strategies (hash-based joins and aggregations) can do a poor job distributing the computation. In this paper, we propose a two-phase execution strategy for numerical computations that are expressed relationally, as aggregated join trees (that is, expressed as a series of relational joins followed by an aggregation). In a pilot run, lineage information is collected; this lineage is used to optimally plan the computation at the level of individual records. Then, the computation is actually executed. We show experimentally that a relational system making use of this two-phase strategy can be an excellent platform for distributed ML computations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- DISTAL: the distributed tensor algebra compilerRohan Yadav, Alex Aiken, Fredrik KjolstadPLDI 2022 · 29 citations
- JoinBoost: Grow Trees Over Normalized Data Using Only SQLZezhou Huang, Rathijit Sen, Jiaxiang Liu, Eugene WuVLDB 2023 · 23 citations
- In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data ShuffleLijie Xu, Shuang Qiu, Binhang Yuan, Jiawei Jiang et al.SIGMOD 2022 · 11 citations
- Auto-Differentiation of Relational Computations for Very Large Scale Machine LearningYuxin Tang, Zhimin Ding, Dimitrije Jankov, Binhang Yuan et al.ICML 2023 · 7 citations
- Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor ComputationYuxin Tang, Zhiyuan Xin, Zhimin Ding, Xinyu Yao et al.VLDB 2026
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Memory-Efficient Pipeline-Parallel DNN TrainingDeepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen et al.ICML 2021 · 283 citations
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm et al.SIGMOD 2020 · 27 citations
Related papers
- Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear AlgebraShangyu Luo, Dimitrije Jankov, Binhang Yuan, Chris JermaineSIGMOD 2021 · 9 citations
- A Practical Approach to Groupjoin and Nested AggregatesPhilipp Fent, Thomas NeumannVLDB 2021 · 11 citations
- Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL PredicatesMingxi Liu, Zhengyuan Ding, Chenyang Zhang, Qingfeng Pan et al.SIGMOD 2026
- T3: Accurate and Fast Performance Prediction for Relational Database Systems With Compiled Decision TreesMaximilian Rieger, Thomas NeumannSIGMOD 2025 · 5 citations
- Lightweight Materialization for Fast Dashboards Over JoinsZezhou Huang, Eugene WuSIGMOD 2024 · 2 citations
