TANGOS: Regularizing Tabular Neural Networks through Gradient Orthogonalization and Specialization
Alan Jeffares, Tennison Liu, Jonathan Crabbé, Fergus Imrie, Mihaela van der Schaar
Abstract
Despite their success with unstructured data, deep neural networks are not yet a panacea for structured tabular data. In the tabular domain, their efficiency crucially relies on various forms of regularization to prevent overfitting and provide strong generalization performance. Existing regularization techniques include broad modelling decisions such as choice of architecture, loss functions, and optimization methods. In this work, we introduce Tabular Neural Gradient Orthogonalization and Specialization (TANGOS), a novel framework for regularization in the tabular setting built on latent unit attributions. The gradient attribution of an activation with respect to a given input feature suggests how the neuron attends to that feature, and is often employed to interpret the predictions of deep networks. In TANGOS, we take a different approach and incorporate neuron attributions directly into training to encourage orthogonalization and specialization of latent attributions in a fully-connected network. Our regularizer encourages neurons to focus on sparse, non-overlapping input features and results in a set of diverse and specialized latent units. In the tabular domain, we demonstrate that our approach can lead to improved out-of-sample generalization performance, outperforming other popular regularization methods. We provide insight into why our regularizer is effective and demonstrate that TANGOS can be applied jointly with existing methods to achieve even greater generalization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55707cce-b78b-4b31-b6d8-a204c87eed16Cited by top-tier papers15
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 52 citations
- Joint Training of Deep Ensembles Fails Due to Learner CollusionAlan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 34 citations
- BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield ModelChenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li et al.ICML 2024 · 26 citations
- High dimensional, tabular deep learning with an auxiliary knowledge graphCamilo Ruiz, Hongyu Ren, Kexin Huang, Jure LeskovecNeurIPS 2023 · 26 citations
- Canonical normalizing flows for manifold learningKyriakos Flouris, Ender KonukogluNeurIPS 2023 · 19 citations
Builds on14
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- VIME: Extending the Success of Self- and Semi-supervised Learning to Tabular DomainJinsung Yoon, Yao Zhang, James Jordon, Mihaela van der SchaarNeurIPS 2020 · 370 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- Well-tuned Simple Nets Excel on Tabular DatasetsArlind Kadra, Marius Lindauer, Frank Hutter, Josif GrabockaNeurIPS 2021 · 288 citations
Related papers
- InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature InterpretationJacob Yoke Hong Si, Wendy Yusi Cheng, Michael Cooper, Rahul G. KrishnanICML 2024 · 15 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Understanding and Enforcing Weight Disentanglement in Task ArithmeticShangge Liu, Yuehan Yin, Lei Wang, Qi Fan et al.CVPR 2026 · 3 citations
- Task-Agnostic Undesirable Feature Deactivation Using Out-of-Distribution DataDongmin Park, Hwanjun Song, Minseok Kim, Jae-Gil LeeNeurIPS 2021 · 8 citations
- Sparse tree-based Initialization for Neural NetworksPatrick Lutz, Ludovic Arnould, Claire Boyer, Erwan ScornetICLR 2023
