xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
Daniel Beaglehole, David Holzmüller, Adityanarayanan Radhakrishnan, Mikhail Belkin
Abstract
Inference from tabular data, collections of continuous and categorical variables organized into matrices, is a foundation for modern technology and science. Yet, in contrast to the dramatic developments in the rest of AI, the best practice for these predictive tasks has been relatively unchanged and is still primarily based on variations of Gradient Boosted Decision Trees (GBDTs). Very recently, there has been renewed interest in developing state-of-the-art methods for tabular data based on recent developments in neural networks and feature learning methods. In this work, we introduce xRFM, an algorithm that combines feature learning kernel machines with a tree structure to both adapt to the local structure of the data and scale to essentially unlimited amounts of training data. On the TALENT benchmark, we show that compared to 31 other methods, including recently introduced tabular foundation models (TabPFNv2) and GBDTs, xRFM achieves best performance across 100 regression datasets and is competitive to the best methods across 200 classification datasets outperforming GBDTs. Additionally, xRFM provides interpretability natively through the Average Gradient Outer Product. Code for xRFM (following a scikit-learn-style API) is available at: https://github.com/dmbeaglehole/xRFM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 85 citations
- The Golden Subspace: Where Efficiency Meets Generalization in Continual Test-Time AdaptationGuannan Lai, Da-Wei Zhou, Zhenguo Li, Han-Jia YeCVPR 2026 · 2 citations
- SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early ExitSi-Yang Liu, Han-Jia YeICML 2026
Builds on8
- On Embeddings for Numerical Features in Tabular Deep LearningYury Gorishniy, Ivan Rubachev, Artem BabenkoNeurIPS 2022 · 338 citations
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- Better by default: Strong pre-tuned MLPs and boosted trees on tabular dataDavid Holzmüller, Léo Grinsztajn, Ingo SteinwartNeurIPS 2024 · 141 citations
- Average gradient outer product as a mechanism for deep neural collapseDaniel Beaglehole, Peter Súkeník, Marco Mondelli, Mikhail BelkinNeurIPS 2024 · 26 citations
- Toward Large Kernel ModelsAmirhesam Abedsoltan, Mikhail Belkin, Parthe PanditICML 2023 · 23 citations
Related papers
- iLTM: Integrated Large Tabular ModelDavid Bonet, Marçal Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat et al.KDD 2026 · 4 citations
- TabICL: A Tabular Foundation Model for In-Context Learning on Large DataJingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le MorvanICML 2025
- Revisiting Nearest Neighbor for Tabular Data: A Deep Tabular Baseline Two Decades LaterHan-Jia Ye, Huai-Hong Yin, De-Chuan Zhan, Wei-Lun ChaoICLR 2025
- NRGBoost: Energy-Based Generative Boosted TreesJoão BravoICLR 2025
- TabR: Tabular Deep Learning Meets Nearest NeighborsYury Gorishniy, Ivan Rubachev, Nikolay Kartashev, Daniil Shlenskii et al.ICLR 2024 · 78 citations
