APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
Hong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih Peng
摘要
Tabular data are fundamental in common machine learning applications, ranging from finance to genomics and healthcare. This paper focuses on tabular regression tasks, a field where deep learning (DL) methods are not consistently superior to machine learning (ML) models due to the challenges posed by irregular target functions inherent in tabular data, causing sensitive label changes with minor variations from features. To address these issues, we propose a novel Arithmetic-Aware Pre-training and Adaptive-Regularized Fine-tuning framework (APAR), which enables the model to fit irregular target function in tabular data while reducing the negative impact of overfitting. In the pre-training phase, APAR introduces an arithmeticaware pretext objective to capture intricate sample-wise relationships from the perspective of continuous labels. In the fine-tuning phase, a consistency-based adaptive regularization technique is proposed to self-learn appropriate data augmentation. Extensive experiments across 10 datasets demonstrated that APAR outperforms existing GBDT-, supervised NN-, and pretrain-finetune NN-based methods in RMSE (+9.43% ∼ 20.37%), and empirically validated the effects of pre-training tasks, including the study of arithmetic operations. Our code and data are publicly available at https://github.com/johnnyhwu/APAR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep LearningJannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez 等NeurIPS 2021 · 被引用 180 次
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii 等ICLR 2024 · 被引用 78 次
- An Inductive Bias for Tabular Deep LearningEge Beyazit, Jonathan Kozaczuk, Bo Li, Vanessa Wallace 等NeurIPS 2023 · 被引用 26 次
相关 Paper
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
- Arithmetic Feature Interaction Is Necessary for Deep Tabular LearningYi Cheng, Renjun Hu, Haochao Ying, Xing Shi 等AAAI 2024 · 被引用 18 次
- XTab: Cross-table Pretraining for Tabular TransformersBingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li 等ICML 2023 · 被引用 111 次
- Generative Table Pre-training Empowers Models for Tabular PredictionTianping Zhang, Shaowen Wang, Shuicheng Yan, Li Jian 等EMNLP 2023 · 被引用 18 次
- Efficient Piecewise-Linear Embeddings for Deep Tabular Regression by Guided Breakpoint AllocationMin-Kook Suh, Moonjung Eo, Kyungeun Lee, Seoyoon Kim 等KDD 2026
