Unveiling the Role of Data Uncertainty in Tabular Deep Learning
Nikolay Kartashev, Ivan Rubachev, Artem Babenko
Abstract
Recent advancements in tabular deep learning have demonstrated exceptional practical performance, yet the field often lacks a clear understanding of why these techniques actually succeed. To address this gap, our paper highlights the importance of the concept of data (aleatoric) uncertainty for explaining the effectiveness of recent tabular DL methods. While data uncertainty leads to irreducible prediction errors on test samples, it also introduces stochasticity into the training signal that can impede effective learning. We demonstrate that tabular methods differ significantly in their ability to cope with this optimization challenge. Specifically, we reveal that the success of many beneficial design choices in tabular DL, such as numerical feature embeddings, advanced ensembling strategies, retrieval-augmented models, and tabular Prior-Fitted Networks, can be partially attributed to their respective implicit mechanisms for performing well under high data uncertainty. By dissecting these varied mechanisms, we provide a unifying understanding of recent performance improvements. Furthermore, leveraging insights from this perspective, we design a novel, more effective numerical feature embedding method as an immediate practical outcome of our analysis. Overall, our work paves the way toward a principled understanding of the benefits introduced by modern tabular methods that results in the concrete advancements of existing techniques and outlines future research directions for tabular DL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- NGBoost: Natural Gradient Boosting for Probabilistic PredictionTony Duan, Anand Avati, Daisy Yi Ding, Khanh K. Thai et al.ICML 2020 · 433 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- Scarf: Self-Supervised Contrastive Learning using Random Feature CorruptionDara Bahri, Heinrich Jiang, Yi Tay, Donald MetzlerICLR 2022 · 233 citations
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
Related papers
- TabM: Advancing tabular deep learning with parameter-efficient ensemblingYury Gorishniy, Akim Kotelnikov, Artem BabenkoICLR 2025
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii et al.ICLR 2024 · 78 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Efficient Piecewise-Linear Embeddings for Deep Tabular Regression by Guided Breakpoint AllocationMin-Kook Suh, Moonjung Eo, Kyungeun Lee, Seoyoon Kim et al.KDD 2026
- Dense Representation Learning and Retrieval for Tabular Data PredictionLei Zheng, Ning Li, Xianyu Chen, Quan Gan et al.KDD 2023 · 7 citations
