A faster training algorithm for regression trees with linear leaves, and an analysis of its complexity
Kuat Gazizov, Miguel Á. Carreira-Perpiñán
Abstract
We consider the Tree Alternating Optimization (TAO) algorithm to train regression trees with linear predictors in the leaves. Unlike the traditional, greedy recursive partitioning algorithms such as CART, TAO guarantees a monotonic decrease of the objective function and results in smaller trees of much better accuracy. We modify the TAO algorithm so that it produces exactly the same result but is much faster, particularly for high input dimensionality or deep trees. The idea is based on the fact that, at each iteration of TAO, each leaf receives only a subset of the training instances. Thus, the optimization of the leaf model can be done exactly but faster by using the Sherman-Morrison-Woodbury formula. This has the unexpected advantage that, once a tree exceeds a critical depth, then making it deeper makes it faster to train, even though the tree is larger and has more parameters. Indeed, this can make learning a nonlinear model (the tree) asymptotically faster than a regular linear regression model. We analyze the corresponding computational complexity and verify the speedups experimentally in various datasets. The argument can be applied to other types of trees, whenever the optimization of a node can be computed in superlinear time of the number of instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2a55d32-0347-473e-9c4f-b0d057c96849Builds on7
- Smaller, more accurate regression forests using tree alternating optimizationArman Zharmagambetov, Miguel Á. Carreira-PerpiñánICML 2020 · 34 citations
- Optimal Interpretable Clustering Using Oblique Decision TreesMagzhan Gabidolla, Miguel Á. Carreira-PerpiñánKDD 2022 · 16 citations
- Pushing the Envelope of Gradient Boosting Forests via Globally-Optimized Oblique TreesMagzhan Gabidolla, Miguel Á. Carreira-PerpiñánCVPR 2022 · 13 citations
- Piecewise Constant and Linear Regression Trees: An Optimal Dynamic Programming ApproachMim van den Bos, Jacobus G. M. van der Linden, Emir DemirovicICML 2024 · 6 citations
- The tree autoencoder model, with application to hierarchical data visualizationMiguel Á. Carreira-Perpiñán, Kuat GazizovNeurIPS 2024 · 3 citations
Related papers
- Softmax Tree: An Accurate, Fast Classifier When the Number of Classes Is LargeArman Zharmagambetov, Magzhan Gabidolla, Miguel Á. Carreira-PerpiñánEMNLP 2021 · 4 citations
- Breiman meets Bellman: Non-Greedy Decision Trees with MDPsHector Kohler, Riad Akrour, Philippe PreuxKDD 2025
- MABSplit: Faster Forest Training Using Multi-Armed BanditsMo Tiwari, Ryan Kang, Jaeyong Lee, Chris Piech et al.NeurIPS 2022 · 5 citations
- Harnessing the power of choices in decision tree learningGuy Blanc, Jane Lange, Chirag Pabbaraju, Colin Sullivan et al.NeurIPS 2023 · 3 citations
- Sparse Learning with CARTJason M. KlusowskiNeurIPS 2020 · 31 citations
