Can a Deep Learning Model be a Sure Bet for Tabular Prediction?
Jintai Chen, Jiahuan Yan, Qiyuan Chen, Danny Z. Chen, Jian Wu, Jimeng Sun
摘要
Data organized in tabular format is ubiquitous in real-world applications, and users often craft tables with biased feature definitions and flexibly set prediction targets of their interests. Thus, a rapid development of a robust, effective, dataset-versatile, user-friendly tabular prediction approach is highly desired. While Gradient Boosting Decision Trees (GBDTs) and existing deep neural networks (DNNs) have been extensively utilized by professional users, they present several challenges for casual users, particularly: (i) the dilemma of model selection due to their different dataset preferences, and (ii) the need for heavy hyperparameter searching, failing which their performances are deemed inadequate. In this paper, we delve into this question: Can we develop a deep learning model that serves as a sure bet solution for a wide range of tabular prediction tasks, while also being user-friendly for casual users? We delve into three key drawbacks of deep tabular models, encompassing: (P1) lack of rotational variance property, (P2) large data demand, and (P3) over-smooth solution. We propose ExcelFormer, addressing these challenges through a semi-permeable attention module that effectively constrains the influence of less informative features to break the DNNs' rotational invariance property (for P1), data augmentation approaches tailored for tabular data (for P2), and attentive feedforward network to boost the model fitting capability (for P3). These designs collectively make ExcelFormer a sure bet solution for diverse tabular datasets. Extensive and stratified experiments conducted on real-world datasets demonstrate that our model outperforms previous approaches across diverse tabular data prediction tasks, and this framework can be friendly to casual users, offering ease of use without the heavy hyperparameter tuning. The codes are available at https://github.com/whatashot/excelformer.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper10
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 被引用 141 次
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 被引用 52 次
- Effective Intra-Inter Interaction Learning for Relational TablesWeichen Li, Ken Zhong, Zheng Wang, Li Pan 等KDD 2026
- Small Models are LLM Knowledge Triggers for Medical Tabular PredictionJiahuan Yan, Jintai Chen, Chaowen Hu, Bo Zheng 等ICLR 2025
- Ripple Perturbations Through Structure: Likelihood-Constrained Adversarial Attacks on Heterogeneous Tabular DataZhengjie Zhou, Jiahuan Yan, Boqun Ma, Weiwei Feng 等ICML 2026
相关 Paper
- Arithmetic Feature Interaction Is Necessary for Deep Tabular LearningYi Cheng, Renjun Hu, Haochao Ying, Xing Shi 等AAAI 2024 · 被引用 18 次
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
- Team up GBDTs and DNNs: Advancing Efficient and Effective Tabular Prediction with Tree-hybrid MLPsJiahuan Yan, Jintai Chen, Qianxing Wang, Danny Z. Chen 等KDD 2024 · 被引用 9 次
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-TuningHong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih PengAAAI 2025 · 被引用 1 次
- T2G-FORMER: Organizing Tabular Features into Relation Graphs Promotes Heterogeneous Feature InteractionJiahuan Yan, Jintai Chen, Yixuan Wu, Danny Z. Chen 等AAAI 2023 · 被引用 60 次
