TabFlex: Scaling Tabular Learning to Millions with Linear Attention
Yuchen Zeng, Tuan Dinh, Wonjun Kang, Andreas C. Mueller
Abstract
Recent advances in the field of in-context learning (ICL) have demonstrated impressive performance for tabular classification, exemplified by TABPFN's success on small datasets. However, the quadratic complexity of the attention mechanism limits its applicability to larger datasets. To address this issue, we conduct a comprehensive comparison of popular scalable attention alternatives, including state-space models (SSMs) and linear attention mechanisms, revealing that the inherent causality of SSMs hinders ICL performance for large datasets, while linear attention preserves effectiveness. Leveraging these insights, we introduce TABFLEX, a model based on linear attention that supports thousands of features and hundreds of classes, capable of handling datasets with millions of samples. Extensive experiments demonstrate that TABFLEX is significantly faster than most existing methods while achieving top-two performance on small datasets among 25 baselines, with a 2× speedup over TABPFN and a 1.5× speedup over XGBoost. On large datasets, TABFLEX remains efficient (e.g., approximately 5 seconds on the poker-hand dataset, which consists of millions of samples), while achieving relatively solid performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0bbd5029-7fd0-44fe-bab8-9aff8da63154Cited by top-tier papers5
- ICR-RL: Deep Reinforcement Learning via In-Context-RegressionDavid Schiff, Ofir Lindenbaum, Yonathan EfroniICML 2026 · 5 citations
- When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment ApproachXinpeng Lv, Yunxin Mao, Renzhe Xu, Chunyuan Zheng et al.ICML 2026 · 2 citations
- EquiTabPFN: A Target-Permutation Equivariant Prior Fitted NetworkMichael Arbel, David Salinas, Frank HutterNeurIPS 2025 · 1 citation
- Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation modelsElias Chaibub NetoICLR 2026
- Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior AdjustmentSeunghan LeeICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- TabICL: A Tabular Foundation Model for In-Context Learning on Large DataJingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le MorvanICML 2025
- End-to-End Compression for Tabular Foundation ModelsGuri Zabërgja, Rafiq Kamel, Arlind Kadra, Christian Frey et al.ICML 2026 · 4 citations
- Mixture of In-Context Prompters for Tabular PFNsDerek Qiang Xu, F. Olcay Cirit, Reza Asadi, Yizhou Sun et al.ICLR 2025
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 96 citations
- SwiftPFN: Revisiting Row-Wise Attention–Only Tabular Foundation Models with Adaptive Early ExitSi-Yang Liu, Han-Jia YeICML 2026
