Can a Deep Learning Model be a Sure Bet for Tabular Prediction?
Jintai Chen, Jiahuan Yan, Qiyuan Chen, Danny Z. Chen, Jian Wu, Jimeng Sun
Abstract
Data organized in tabular format is ubiquitous in real-world applications, and users often craft tables with biased feature definitions and flexibly set prediction targets of their interests. Thus, a rapid development of a robust, effective, dataset-versatile, user-friendly tabular prediction approach is highly desired. While Gradient Boosting Decision Trees (GBDTs) and existing deep neural networks (DNNs) have been extensively utilized by professional users, they present several challenges for casual users, particularly: (i) the dilemma of model selection due to their different dataset preferences, and (ii) the need for heavy hyperparameter searching, failing which their performances are deemed inadequate. In this paper, we delve into this question: Can we develop a deep learning model that serves as a sure bet solution for a wide range of tabular prediction tasks, while also being user-friendly for casual users? We delve into three key drawbacks of deep tabular models, encompassing: (P1) lack of rotational variance property, (P2) large data demand, and (P3) over-smooth solution. We propose ExcelFormer, addressing these challenges through a semi-permeable attention module that effectively constrains the influence of less informative features to break the DNNs' rotational invariance property (for P1), data augmentation approaches tailored for tabular data (for P2), and attentive feedforward network to boost the model fitting capability (for P3). These designs collectively make ExcelFormer a sure bet solution for diverse tabular datasets. Extensive and stratified experiments conducted on real-world datasets demonstrate that our model outperforms previous approaches across diverse tabular data prediction tasks, and this framework can be friendly to casual users, offering ease of use without the heavy hyperparameter tuning. The codes are available at https://github.com/whatashot/excelformer.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0bac5a6e-4589-4191-9fca-fe7a6e66af8cCited by top-tier papers10
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 52 citations
- Effective Intra-Inter Interaction Learning for Relational TablesWeichen Li, Ken Zhong, Zheng Wang, Li Pan et al.KDD 2026
- Small Models are LLM Knowledge Triggers for Medical Tabular PredictionJiahuan Yan, Jintai Chen, Chaowen Hu, Bo Zheng et al.ICLR 2025
- Ripple Perturbations Through Structure: Likelihood-Constrained Adversarial Attacks on Heterogeneous Tabular DataZhengjie Zhou, Jiahuan Yan, Boqun Ma, Weiwei Feng et al.ICML 2026
Related papers
- Arithmetic Feature Interaction Is Necessary for Deep Tabular LearningYi Cheng, Renjun Hu, Haochao Ying, Xing Shi et al.AAAI 2024 · 18 citations
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
- Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPsJiahuan Yan, Jintai Chen, Qianxing Wang, Danny Z. Chen et al.KDD 2024 · 9 citations
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-TuningHong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih PengAAAI 2025 · 1 citation
- T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature InteractionJiahuan Yan, Jintai Chen, Yixuan Wu, Danny Z. Chen et al.AAAI 2023 · 60 citations
