TabR: Tabular Deep Learning Meets Nearest Neighbors
Yury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii, Akim Kotelnikov, Artem Babenko
Abstract
Deep learning (DL) models for tabular data problems (e.g. classification, regression) are currently receiving increasingly more attention from researchers. However, despite the recent efforts, the non-DL algorithms based on gradient-boosted decision trees (GBDT) remain a strong go-to solution for these problems. One of the research directions aimed at improving the position of tabular DL involves designing so-called retrieval-augmented models. For a target object, such models retrieve other objects (e.g. the nearest neighbors) from the available training data and use their features and labels to make a better prediction. In this work, we present TabR -essentially, a feed-forward network with a custom k-Nearest-Neighbors-like component in the middle. On a set of public benchmarks with datasets up to several million objects, TabR marks a big step forward for tabular DL: it demonstrates the best average performance among tabular DL models, becomes the new state-of-the-art on several datasets, and even outperforms GBDT models on the recently proposed "GBDT-friendly" benchmark (see Figure 1 ). Among the important findings and technical details powering TabR, the main ones lie in the attention-like mechanism that is responsible for retrieving the nearest neighbors and extracting valuable signal from them. In addition to the much higher performance, TabR is simple and significantly more efficient compared to prior retrieval-based tabular DL models. The source code is published: link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3328febd-aea4-4452-9564-ca08e2e39669Cited by top-tier papers16
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach et al.NeurIPS 2025 · 118 citations
- Retrieval & Fine-Tuning for In-Context Tabular ModelsValentin Thomas, Junwei Ma, Rasa Hosseinzadeh, Keyvan Golestan et al.NeurIPS 2024 · 75 citations
- ConTextTab: A Semantics-Aware Tabular In-Context LearnerMarco Spinaci, Marek Polewczyk, Maximilian Schambach, Sam ThelinNeurIPS 2025 · 36 citations
- EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language ModelsJinhee Kim, Taesung Kim, Jaegul ChooNeurIPS 2024 · 22 citations
Builds on17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
Related papers
- TabM: Advancing tabular deep learning with parameter-efficient ensemblingYury Gorishniy, Akim Kotelnikov, Artem BabenkoICLR 2025
- Dense Representation Learning and Retrieval for Tabular Data PredictionLei Zheng, Ning Li, Xianyu Chen, Quan Gan et al.KDD 2023 · 7 citations
- iLTM: Integrated Large Tabular ModelDavid Bonet, Marçal Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat et al.KDD 2026 · 4 citations
- Revisiting Nearest Neighbor for Tabular Data: A Deep Tabular Baseline Two Decades LaterHan-Jia Ye, Huai-Hong Yin, De-Chuan Zhan, Wei-Lun ChaoICLR 2025
- DOFEN: Deep Oblivious Forest ENsembleKuan-Yu Chen, Ping-Han Chiang, Hsin-Rung Chou, Chih-Sheng Chen et al.NeurIPS 2024 · 6 citations
