JoinBoost: Grow Trees Over Normalized Data Using Only SQL
Zezhou Huang, Rathijit Sen, Jiaxiang Liu, Eugene Wu
Abstract
Although dominant for tabular data, ML libraries that train tree models over normalized databases (e.g., LightGBM, XGBoost) require the data to be denormalized as a single table, materialized, and exported. This process is not scalable, slow, and poses security risks. In-DB ML aims to train models within DBMSes to avoid data movement and provide data governance. Rather than modify a DBMS to support In-DB ML, is it possible to offer competitive tree training performance to specialized ML libraries...with only SQL?
We present JoinBoost, a Python library that rewrites tree training algorithms over normalized databases into pure SQL. It is portable to any DBMS, offers performance competitive with specialized ML libraries, and scales with the underlying DBMS capabilities. JoinBoost extends prior work from both algorithmic and systems perspectives. Algorithmically, we support factorized gradient boosting, by updating the Y variable to the residual in the non-materialized join result. Although this view update problem is generally ambiguous, we identify addition-to-multiplication preserving , the key property of variance semi-ring to support rmse the most widely used criterion. System-wise, we identify residual updates as a performance bottleneck. Such overhead can be natively minimized on columnar DBMSes by creating a new column of residual values and adding it as a projection. We validate this with two implementations on DuckDB, with no or minimal modifications to its internals for portability. Our experiment shows that JoinBoost is 3× (1.1×) faster for random forests (gradient boosting) compared to LightGBM, and over an order of magnitude faster than state-of-the-art In-DB ML systems. Further, JoinBoost scales well beyond LightGBM in terms of the # features, DB size (TPC-DS SF=1000), and join graph complexity (galaxy schemas).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Saibot: A Differentially Private Data Search PlatformZezhou Huang, Jiaxiang Liu, Daniel Alabi, Raul Castro Fernandez et al.VLDB 2023 · 13 citations
- Quantum Data Management in the NISQ EraRihan Hai, Shih-Han Hung, Tim Coopmans, Tim Littau et al.VLDB 2025 · 10 citations
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 5 citations
- Biathlon: Harnessing Model Resilience for Accelerating ML Inference PipelinesChaokun Chang, Eric Lo, Chunxiao YeVLDB 2024 · 5 citations
- InferF: Declarative Factorization of AI/ML Inferences over JoinsKanchan Chowdhury, Lixi Zhou, Lulu Xie, Xinwei Fu et al.SIGMOD 2026 · 2 citations
Builds on5
- TCUDB: Accelerating Database with Tensor ProcessorsYu-Ching Hu, Yuliang Li, Hung-Wei TsengSIGMOD 2022 · 32 citations
- ARDA: Automatic Relational Data Augmentation for Machine LearningNadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez et al.VLDB 2020 · 14 citations
- PGMJoins: Random Join Sampling with Graphical ModelsAli Mohammadi Shanghooshabad, Meghdad Kurmanji, Qingzhi Ma, Michael Shekelyan et al.SIGMOD 2021 · 13 citations
- Towards Factorized SVM with Gaussian Kernels over Normalized DataKeyu Yang, Yunjun Gao, Lei Liang, Bin Yao et al.ICDE 2020 · 13 citations
- Distributed Numerical and Machine Learning Computations via Two-Phase Execution of Aggregated Join TreesDimitrije Jankov, Binhang Yuan, Shangyu Luo, Chris JermaineVLDB 2021 · 9 citations
Related papers
- Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL PredicatesMingxi Liu, Zhengyuan Ding, Chenyang Zhang, Qingfeng Pan et al.SIGMOD 2026
- InferDB: In-Database Machine Learning Inference Using IndexesRicardo Salazar-Díaz, Boris Glavic, Tilmann RablVLDB 2024 · 13 citations
- DistVec: Efficient Distributed Machine Learning in Parallel Database SystemsXinyi Zhang, Liangzu Liu, Xupeng Miao, Yinjun Wu et al.ICDE 2026
- Machine Learning Inference Pipeline Execution Using Pure SQL Based on Operator FusionQingfeng Pan, Jiahe Zhi, Chenyang Zhang, Chen Xu et al.ICDE 2025 · 2 citations
- DB4ML - An In-Memory Database Kernel with Machine Learning SupportMatthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Röhm et al.SIGMOD 2020 · 27 citations
