Retrieval & Fine-Tuning for In-Context Tabular Models
Valentin Thomas, Junwei Ma, Rasa Hosseinzadeh, Keyvan Golestan, Guangwei Yu, Maksims Volkovs, Anthony L. Caterini
摘要
Tabular data is a pervasive modality spanning a wide range of domains, and the inherent diversity poses a considerable challenge for deep learning. Recent advancements using transformer-based in-context learning have shown promise on smaller and less complex datasets, but have struggled to scale to larger and more complex ones. To address this limitation, we propose a combination of retrieval and fine-tuning: we can adapt the transformer to a local subset of the data by collecting nearest neighbours, and then perform task-specific fine-tuning with this retrieved set of neighbours in context. Using TabPFN as the base model -- currently the best tabular in-context learner -- and applying our retrieval and fine-tuning scheme on top results in what we call a locally-calibrated PFN, or LoCalPFN. We conduct extensive evaluation on 95 datasets curated by TabZilla from OpenML, upon which we establish a new state-of-the-art with LoCalPFN -- even with respect to tuned tree-based models. Notably, we show a significant boost in performance compared to the base in-context model, demonstrating the efficacy of our approach and advancing the frontier of deep learning in tabular data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 被引用 85 次
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its CapabilitiesHan-Jia Ye, Si-Yang Liu, Wei-Lun ChaoNeurIPS 2025 · 被引用 52 次
- CausalPFN: Amortized Causal Effect Estimation via In-Context LearningVahid Balazadeh Meresht, Hamidreza Kamkari, Valentin Thomas, Junwei Ma 等NeurIPS 2025 · 被引用 52 次
- Effortless, Simulation-Efficient Bayesian Inference using Tabular Foundation ModelsJulius Vetter, Manuel Glöckler, Daniel Gedon, Jakob H. MackeNeurIPS 2025 · 被引用 14 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
相关 Paper
- Mixture of In-Context Prompters for Tabular PFNsDerek Qiang Xu, F. Olcay Cirit, Reza Asadi, Yizhou Sun 等ICLR 2025
- TabPFN Unleashed: A Scalable and Effective Solution to Tabular Classification ProblemsSiyang Liu, Han-Jia YeICML 2025
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
- Mitigating Label Shift in Tabular In-Context Learning via Test-Time Posterior AdjustmentSeunghan LeeICML 2026
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde 等NeurIPS 2024 · 被引用 57 次
