APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
Hong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih Peng
Abstract
Tabular data are fundamental in common machine learning applications, ranging from finance to genomics and healthcare. This paper focuses on tabular regression tasks, a field where deep learning (DL) methods are not consistently superior to machine learning (ML) models due to the challenges posed by irregular target functions inherent in tabular data, causing sensitive label changes with minor variations from features. To address these issues, we propose a novel Arithmetic-Aware Pre-training and Adaptive-Regularized Fine-tuning framework (APAR), which enables the model to fit irregular target function in tabular data while reducing the negative impact of overfitting. In the pre-training phase, APAR introduces an arithmeticaware pretext objective to capture intricate sample-wise relationships from the perspective of continuous labels. In the fine-tuning phase, a consistency-based adaptive regularization technique is proposed to self-learn appropriate data augmentation. Extensive experiments across 10 datasets demonstrated that APAR outperforms existing GBDT-, supervised NN-, and pretrain-finetune NN-based methods in RMSE (+9.43% ∼ 20.37%), and empirically validated the effects of pre-training tasks, including the study of arithmetic operations. Our code and data are publicly available at https://github.com/johnnyhwu/APAR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep LearningJannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez et al.NeurIPS 2021 · 180 citations
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii et al.ICLR 2024 · 78 citations
- An Inductive Bias for Tabular Deep LearningEge Beyazit, Jonathan Kozaczuk, Bo Li, Vanessa Wallace et al.NeurIPS 2023 · 26 citations
Related papers
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- Arithmetic Feature Interaction Is Necessary for Deep Tabular LearningYi Cheng, Renjun Hu, Haochao Ying, Xing Shi et al.AAAI 2024 · 18 citations
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li et al.ICML 2023 · 111 citations
- Generative Table Pre-training Empowers Models for Tabular PredictionTianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian et al.EMNLP 2023 · 18 citations
- Efficient Piecewise-Linear Embeddings for Deep Tabular Regression by Guided Breakpoint AllocationMin-Kook Suh, Moonjung Eo, Kyungeun Lee, Seoyoon Kim et al.KDD 2026
