Distributed Task-Based Training of Tree Models
Da Yan, Md Mashiur Rahman Chowdhury, Guimu Guo, Jalal Khalil, Zhe Jiang, Sushil K. Prasad
摘要
Decision trees and tree ensembles are popular supervised learning models on tabular data. Two recent research trends on tree models stand out: (1) bigger and deeper models with many trees, and (2) scalable distributed training frameworks. However, existing implementations on distributed systems are IO-bound leaving CPU cores underutilized. They also only find best node-splitting conditions approximately due to row-based data partitioning scheme. In this paper, we target the exact training of tree models by effectively utilizing the available CPU cores. The resulting system called TreeServer adopts a column-based data partitioning scheme to minimize communication, and a node-centric task-based engine to fully explore the CPU parallelism. Experiments show that TreeServer is up to 10× faster than models in Spark MLlib. We also showcase TreeServer's high training throughput by using it to build big “deep forest” models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
- Treebeard: An Optimizing Compiler for Decision Tree Based ML InferenceAshwin Prasad, Sampath Rajendra, Kaushik Rajan, R. Govindarajan 等MICRO 2022 · 被引用 8 次
- GRANDE: Gradient-Based Decision Tree Ensembles for Tabular DataSascha Marton, Stefan Lüdtke, Christian Bartelt, Heiner StuckenschmidtICLR 2024 · 被引用 13 次
- C olumnSGD: A Column-oriented Framework for Distributed Stochastic Gradient DescentZhipeng Zhang, Wentao Wu, Jiawei Jiang, Lele Yu 等ICDE 2020 · 被引用 6 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
