Retrieval & Fine-Tuning for In-Context Tabular Models
Valentin Thomas, Junwei Ma, Rasa Hosseinzadeh, Keyvan Golestan, Guangwei Yu, Maksims Volkovs, Anthony L. Caterini
Abstract
Tabular data is a pervasive modality spanning a wide range of domains, and the inherent diversity poses a considerable challenge for deep learning. Recent advancements using transformer-based in-context learning have shown promise on smaller and less complex datasets, but have struggled to scale to larger and more complex ones. To address this limitation, we propose a combination of retrieval and fine-tuning: we can adapt the transformer to a local subset of the data by collecting nearest neighbours, and then perform task-specific fine-tuning with this retrieved set of neighbours in context. Using TabPFN as the base model -- currently the best tabular in-context learner -- and applying our retrieval and fine-tuning scheme on top results in what we call a locally-calibrated PFN, or LoCalPFN. We conduct extensive evaluation on 95 datasets curated by TabZilla from OpenML, upon which we establish a new state-of-the-art with LoCalPFN -- even with respect to tuned tree-based models. Notably, we show a significant boost in performance compared to the base in-context model, demonstrating the efficacy of our approach and advancing the frontier of deep learning in tabular data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8690a749-eddc-48d7-8029-ed8618c1c6b0Cited by top-tier papers15
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach et al.NeurIPS 2025 · 118 citations
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 85 citations
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 52 citations
- CausalPFN: Amortized Causal Effect Estimation via In-Context LearningVahid Balazadeh Meresht, Hamidreza Kamkari, Valentin Thomas, Junwei Ma et al.NeurIPS 2025 · 52 citations
- Effortless, Simulation-Efficient Bayesian Inference using Tabular Foundation ModelsJulius Vetter, Manuel Glöckler, Daniel Gedon, Jakob H. MackeNeurIPS 2025 · 14 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
Related papers
- Mixture of In-Context Prompters for Tabular PFNsDerek Qiang Xu, F. Olcay Cirit, Reza Asadi, Yizhou Sun et al.ICLR 2025
- TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification ProblemsSiyang Liu, Han-Jia YeICML 2025
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 96 citations
- Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior AdjustmentSeunghan LeeICML 2026
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde et al.NeurIPS 2024 · 57 citations
