Well-tuned Simple Nets Excel on Tabular Datasets
Arlind Kadra, Marius Lindauer, Frank Hutter, Josif Grabocka
摘要
Tabular datasets are the last "unconquered castle" for deep learning, with traditional ML methods like Gradient-Boosted Decision Trees still performing strongly even against recent specialized neural architectures. In this paper, we hypothesize that the key to boosting the performance of neural networks lies in rethinking the joint and simultaneous application of a large set of modern regularization techniques. As a result, we propose regularizing plain Multilayer Perceptron (MLP) networks by searching for the optimal combination/cocktail of 13 regularization techniques for each dataset using a joint optimization over the decision on which regularizers to apply and their subsidiary hyperparameters. We empirically assess the impact of these regularization cocktails for MLPs in a large-scale empirical study comprising 40 tabular datasets and demonstrate that (i) well-regularized plain MLPs significantly outperform recent state-of-the-art specialized neural network architectures, and (ii) they even outperform strong traditional ML methods, such as XGBoost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 被引用 338 次
- TransTab: Learning Transferable Tabular Transformers Across TablesZifeng Wang, Jimeng SunNeurIPS 2022 · 被引用 242 次
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 被引用 141 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation ModelsXiyuan Zhang, Danielle Maddix Robinson, Junming Yin, Nick Erickson 等NeurIPS 2025 · 被引用 91 次
它引用的顶会 Paper9
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
- TrivialAugment: Tuning-free Yet State-of-the-Art Data AugmentationSamuel G. Müller, Frank HutterICCV 2021 · 被引用 384 次
相关 Paper
- Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPsJiahuan Yan, Jintai Chen, Qianxing Wang, Danny Z. Chen 等KDD 2024 · 被引用 9 次
- TabPack: Efficient Hyperparameter Ensembles for Tabular Deep LearningYury Gorishniy, Akim Kotelnikov, Ivan Rubachev, Artem BabenkoICML 2026
- Sparse tree-based Initialization for Neural NetworksPatrick Lutz, Ludovic Arnould, Claire Boyer, Erwan ScornetICLR 2023
- TabM: Advancing tabular deep learning with parameter-efficient ensemblingYury Gorishniy, Akim Kotelnikov, Artem BabenkoICLR 2025
- TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and SpecializationAlan Jeffares, Tennison Liu, Jonathan Crabbé, Fergus Imrie 等ICLR 2023 · 被引用 3 次
