An Inductive Bias for Tabular Deep Learning
Ege Beyazit, Jonathan Kozaczuk, Bo Li, Vanessa Wallace, Bilal Fadlallah
摘要
Deep learning methods have achieved state-of-the-art performance in most modeling tasks involving images, text and audio, however, they typically underperform tree-based methods on tabular data. In this paper, we hypothesize that a significant contributor to this performance gap is the interaction between irregular target functions resulting from the heterogeneous nature of tabular feature spaces, and the well-known tendency of neural networks to learn smooth functions. Utilizing tools from spectral analysis, we show that functions described by tabular datasets often have high irregularity, and that they can be smoothed by transformations such as scaling and ranking in order to improve performance. However, because these transformations tend to lose information or negatively impact the loss landscape during optimization, they need to be rigorously fine-tuned for each feature to achieve performance gains. To address these problems, we propose introducing frequency reduction as an inductive bias. We realize this bias as a neural network layer that promotes learning low-frequency representations of the input features, allowing the network to operate in a space where the target function is more regular. Our proposed method introduces less computational complexity than a fully connected layer, while significantly improving neural network performance, and speeding up its convergence on 14 tabular datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Retrieval & Fine-Tuning for In-Context Tabular ModelsValentin Thomas, Junwei Ma, Rasa Hosseinzadeh, Keyvan Golestan 等NeurIPS 2024 · 被引用 75 次
- Interpretable Deep Clustering for Tabular DataJonathan Svirsky, Ofir LindenbaumICML 2024 · 被引用 19 次
- Binning as a Pretext Task: Improving Self-Supervised Learning in Tabular DomainsKyungeun Lee, Ye Seul Sim, Hye-Seung Cho, Moonjung Eo 等ICML 2024 · 被引用 17 次
- Adaptive Width Neural NetworksFederico Errica, Henrik Christiansen, Viktor Zaverkin, Mathias Niepert 等ICLR 2026 · 被引用 6 次
- Hybrid Autoencoders for Tabular Data: Leveraging Model-Based Augmentation in Low-Label SettingsErel Naor, Ofir LindenbaumNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper9
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular DomainJinsung Yoon, Yao Zhang, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 被引用 370 次
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 被引用 338 次
相关 Paper
- Spectral Bias in Practice: The Role of Function Frequency in GeneralizationSara Fridovich-Keil, Raphael Gontijo Lopes, Rebecca RoelofsNeurIPS 2022 · 被引用 61 次
- Deep Frequency Principle Towards Understanding Why Deeper Learning Is FasterZhiqin John Xu, Hanxu ZhouAAAI 2021 · 被引用 67 次
- Sparse tree-based Initialization for Neural NetworksPatrick Lutz, Ludovic Arnould, Claire Boyer, Erwan ScornetICLR 2023
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-TuningHong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih PengAAAI 2025 · 被引用 1 次
- Trompt: Towards a Better Deep Neural Network for Tabular DataKuan-Yu Chen, Ping-Han Chiang, Hsin-Rung Chou, Ting-Wei Chen 等ICML 2023 · 被引用 42 次
