TabularBERT: Binning-Based Self-Supervised Learning for Tabular Representation
Beomjin Park, Seunghwan An, Sungchul Hong, Hosik Choi
Abstract
Tabular data is one of the most fundamental and widely used formats for representing structured information. Classical machine learning algorithms continue to achieve substantial success in extracting predictive patterns and constructing accurate models from structured data; however, representation learning approaches that extend language-model-based methods to the tabular setting have opened new opportunities. Nevertheless, conventional tokenization procedures and token embedding mechanisms are not well-suited to numerical variables, as they fail to preserve key numerical properties, including proximity structure and ordinal relationships. To address this limitation, we propose TabularBERT, a Transformer-based model that discretizes numerical variables via binning-based tokenization and learns representations that account for numerical proximity and ordinal information while capturing conditional dependencies among variables through masked self-supervised pretraining. We empirically demonstrate the effectiveness and interpretability of the proposed approach, highlighting the benefits of language-model-based representation learning in the tabular domain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d842b17c-9163-44db-95f2-7085442a1a57Builds on10
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 407 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
Related papers
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
- Numerical Tuple Extraction from Tables with Pre-trainingQingping Yang, Yixuan Cao, Ping LuoKDD 2022 · 1 citation
- Q-Tab: Quantized Tabular Data GeneratorJulian Wustl, Philipp Haid, Yarema Okhrin, Claudius SchnörrICML 2026
- Binning as a Pretext Task: Improving Self-Supervised Learning in Tabular DomainsKyungeun Lee, Ye Seul Sim, Hye-Seung Cho, Moonjung Eo et al.ICML 2024 · 17 citations
