Using maximal information auxiliary variables to improve synthetic data generation based on TabPFN foundation models
Elias Chaibub Neto
摘要
Synthetic data generation for tabular datasets is shifting toward the use of large, general-purpose foundation models. TabPFN, a state-of-the-art example, uses in-context learning to generate probabilistic predictions conditioned on observed examples in a single forward pass. However, when variables are only weakly associated with others, the model's ability to generate realistic synthetic data deteriorates, as the context examples provide little predictive signal. To address this, we introduce the maximal information auxiliary variable (MIAV) strategy, which increases context information with auxiliary variables constructed by rank-matching random noise variables to real data. We establish theoretical properties of the approach which explain its good performance for weakly associated variables. Additional practical advantages of the MIAV approach include improved computational efficiency and invariance to variable order during the synthetic data generation process. Empirical evaluations, on simulated and real datasets, illustrate how the MIAV strategy improves data generation when compared to direct application of TabPFN, and is competitive against other baselines. To illustrate the generality of the MIAV approach we also present an implementation based on the TabICL model (a more scalable tabular foundation model restricted to classification tasks) for performing synthetic data generation on categorical datasets. Overall, MIAV offers an effective foundation model–based alternative to bespoke synthetic data generators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
相关 Paper
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 被引用 85 次
- TabICL: A Tabular Foundation Model for In-Context Learning on Large DataJingang Qu, David Holzmüller, Gaël Varoquaux, Marine Le MorvanICML 2025
- Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation ModelsXiyuan Zhang, Danielle Maddix Robinson, Junming Yin, Nick Erickson 等NeurIPS 2025 · 被引用 91 次
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
- Mixture of In-Context Prompters for Tabular PFNsDerek Qiang Xu, F. Olcay Cirit, Reza Asadi, Yizhou Sun 等ICLR 2025
