Efficient Piecewise-Linear Embeddings for Deep Tabular Regression by Guided Breakpoint Allocation
Min-Kook Suh, Moonjung Eo, Kyungeun Lee, Seoyoon Kim, Hye-Seung Cho, Woohyung Lim
Abstract
Tabular prediction is central to scientific and industrial decision-making, and many high-impact use cases are regression tasks that predict continuous targets. Deep tabular networks for regression tasks rely on effectively modeling irregular and non-smooth relationships between numerical features and continuous targets. To this end, numerous numerical embeddings have been proposed. However, their performance is often highly sensitive to parameter choices, such as the choice of frequencies for Fourier features and the placement of breakpoints for piecewise-linear embeddings. Moreover, the irregular and non-smooth characteristics of tabular data make it difficult for gradient-based end-to-end training to discover effective parameter settings for these embeddings. In this work, we demonstrate the underexplored potential of piecewise-linear embeddings, showing that finding better breakpoints alone can yield substantial gains. We introduce GBDT-Guided Piecewise-Linear (GGPL) embeddings, which leverage GBDTs to provide a strong data-driven prior for breakpoint placement. To further fine-tune these breakpoints via gradient descent, GGPL reparameterizes them for numerical stability and regularizes training via stochastic breakpoint deactivation. Across 28 regression datasets, integrating GGPL with diverse state-of-the-art deep tabular models yields consistent and significant improvements. GGPL is applied only at training time, introducing no inference overhead. Furthermore, in controlled MLP experiments, it uses only 0.4× as many breakpoints on average as prior piecewise-linear embeddings, while achieving higher accuracy. Combined with its negligible overhead, these results establish GGPL as an effective numerical embedding for deep tabular regression.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- iLTM: Integrated Large Tabular ModelDavid Bonet, Marçal Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat et al.KDD 2026 · 4 citations
- APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-TuningHong-Wei Wu, Wei-Yao Wang, Kuang-Da Wang, Wen-Chih PengAAAI 2025 · 1 citation
- Unveiling the Role of Data Uncertainty in Tabular Deep LearningNikolay Kartashev, Ivan Rubachev, Artem BabenkoICML 2026 · 1 citation
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
