TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning Benchmarks
Ivan Rubachev, Nikolay Kartashev, Yury Gorishniy, Artem Babenko
Abstract
Advances in machine learning research drive progress in real-world applications. To ensure this progress, it is important to understand the potential pitfalls on the way from a novel method's success on academic benchmarks to its practical deployment. In this work, we analyze existing tabular deep learning benchmarks and find two common characteristics of tabular data in typical industrial applications that are underrepresented in the datasets usually used for evaluation in the literature. First, in real-world deployment scenarios, distribution of data often changes over time. To account for this distribution drift, time-based train/test splits should be used in evaluation. However, existing academic tabular datasets often lack timestamp metadata to enable such evaluation. Second, a considerable portion of datasets in production settings stem from extensive data acquisition and feature engineering pipelines. This can have an impact on the absolute and relative number of predictive, uninformative, and correlated features compared to academic datasets. In this work, we aim to understand how recent research advances in tabular deep learning transfer to these underrepresented conditions. To this end, we introduce TabReD -a collection of eight industry-grade tabular datasets. We reassess a large number of tabular ML models and techniques on TabReD. We demonstrate that evaluation on both time-based data splits and richer feature sets leads to different methods ranking, compared to evaluation on random splits and smaller number of features, which are common in academic benchmarks. Furthermore, simple MLP-like architectures and GBDT show the best results on the TabReD datasets, while other methods are less effective in the new setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8eb47b94-29fc-4205-9093-db64942cfb35Cited by top-tier papers7
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- iLTM: Integrated Large Tabular ModelDavid Bonet, Marçal Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat et al.KDD 2026 · 4 citations
- Inverse Entropic Optimal Transport Solves Semi-supervised Learning via Data Likelihood MaximizationMikhail Persiianov, Arip Asadulaev, Nikita Andreev, Nikita Starodubcev et al.ICML 2026 · 2 citations
- LassoFlexNet: a Flexible Neural Architecture for Tabular DataKry Yik Chau Lui, Cheng Chi, Kishore Basu, Yanshuai CaoICML 2026
- Epistemic Uncertainty Quantification To Improve Decisions From Black-Box ModelsSébastien Melo, Gaël Varoquaux, Marine Le MorvanICLR 2026
Builds on15
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
Related papers
- TableBench: A Comprehensive and Complex Benchmark for Table Question AnsweringXianjie Wu, Jian Yang, Linzheng Chai, Ge Zhang et al.AAAI 2025 · 138 citations
- TabFSBench: Tabular Benchmark for Feature Shifts in Open EnvironmentsZi-Jian Cheng, Ziyi Jia, Zhi Zhou, Yufeng Li et al.ICML 2025
- TAB: Unified Benchmarking of Time Series Anomaly Detection MethodsXiangfei Qiu, Zhe Li, Wanghui Qiu, Shiyan Hu et al.VLDB 2025 · 57 citations
- Trompt: Towards a Better Deep Neural Network for Tabular DataKuan-Yu Chen, Ping-Han Chiang, Hsin-Rung Chou, Ting-Wei Chen et al.ICML 2023 · 42 citations
- T2R-BENCH: A Benchmark for Real World Table-to-Report TaskJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei et al.EMNLP 2025 · 2 citations
