Learning Binary Decision Trees by Argmin Differentiation
Valentina Zantedeschi, Matt J. Kusner, Vlad Niculae
摘要
We address the problem of learning binary decision trees that partition data for some downstream task. We propose to learn discrete parameters (i.e., for tree traversals and node pruning) and continuous parameters (i.e., for tree split functions and prediction functions) simultaneously using argmin differentiation. We do so by sparsely relaxing a mixed-integer program for the discrete parameters, to allow gradients to pass through the program to continuous parameters. We derive customized algorithms to efficiently compute the forward and backward passes. This means that our tree learning procedure can be used as an (implicit) layer in arbitrary deep networks, and can be optimized with arbitrary loss functions. We demonstrate that our approach produces binary trees that are competitive with existing single tree and ensemble approaches, in both supervised and unsupervised settings. Further, apart from greedy approaches (which do not have competitive accuracies), our method is faster to train than all other tree-learning baselines we compare with. The code for reproducing the results is available at https://github.com/ vzantedeschi/LatentTrees .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Deep Differentiable Logic Gate NetworksFelix Petersen, Christian Borgelt, Hilde Kuehne, Oliver DeussenNeurIPS 2022 · 被引用 117 次
- GradTree: Learning Axis-Aligned Decision Trees with Gradient DescentSascha Marton, Stefan Lüdtke, Christian Bartelt, Heiner StuckenschmidtAAAI 2024 · 被引用 17 次
- Differentiable Decision Tree via "ReLU+Argmin" ReformulationQiangqiang Mao, Jiayang Ren, Yixiu Wang, Chenxuanyin Zou 等NeurIPS 2025 · 被引用 2 次
- Hierarchical Retrieval at Scale: Bridging Interpretability and EfficiencyShubham Gupta, Zichao Li, Tianyi Chen, Cem Subakan 等ICML 2026
- Deep Networks Learn Features From Local Discontinuities in the Label FunctionPrithaj Banerjee, Harish Guruprasad Ramaswamy, Mahesh Lorik Yadav, Chandra Shekar LakshminarayananICLR 2025
它引用的顶会 Paper5
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
- Generalized and Scalable Optimal Sparse Decision TreesJimmy Lin, Chudi Zhong, Diane Hu, Cynthia Rudin 等ICML 2020 · 被引用 174 次
- The Tree Ensemble Layer: Differentiability meets Conditional ComputationHussein Hazimeh, Natalia Ponomareva, Petros Mol, Zhenyu Tan 等ICML 2020 · 被引用 95 次
- A Scalable MIP-based Method for Learning Optimal Multivariate Decision TreesHaoran Zhu, Pavankumar Murali, Dzung T. Phan, Lam M. Nguyen 等NeurIPS 2020 · 被引用 47 次
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie 等CVPR 2020
相关 Paper
- Quant-BnB: A Scalable Branch-and-Bound Method for Optimal Decision Trees with Continuous FeaturesRahul Mazumder, Xiang Meng, Haoyue WangICML 2022 · 被引用 21 次
- Optimal Classification Trees for Continuous Feature Data Using Dynamic Programming with Branch-and-BoundCatalin E. Brita, Jacobus G. M. van der Linden, Emir DemirovicAAAI 2025 · 被引用 5 次
- Discrete Tree Flows via Tree-Structured PermutationsMai Elkady, Hyung Zin Lim, David I. InouyeICML 2022 · 被引用 2 次
- Feature Learning for Interpretable, Performant Decision TreesJack H. Good, Torin Kovach, Kyle Miller, Artur DubrawskiNeurIPS 2023 · 被引用 16 次
- Flexible Modeling and Multitask Learning using Differentiable Tree EnsemblesShibal Ibrahim, Hussein Hazimeh, Rahul MazumderKDD 2022 · 被引用 3 次
