TabEBM: A Tabular Data Augmentation Method with Distinct Class-Specific Energy-Based Models
Andrei Margeloiu, Xiangjian Jiang, Nikola Simidjievski, Mateja Jamnik
摘要
Data collection is often difficult in critical fields such as medicine, physics, and chemistry. As a result, classification methods usually perform poorly with these small datasets, leading to weak predictive performance. Increasing the training set with additional synthetic data, similar to data augmentation in images, is commonly believed to improve downstream classification performance. However, current tabular generative methods that learn either the joint distribution or the class-conditional distribution often overfit on small datasets, resulting in poor-quality synthetic data, usually worsening classification performance compared to using real data alone. To solve these challenges, we introduce TabEBM, a novel class-conditional generative method using Energy-Based Models (EBMs). Unlike existing methods that use a shared model to approximate all class-conditional densities, our key innovation is to create distinct EBM generative models for each class, each modelling its class-specific data distribution individually. This approach creates robust energy landscapes, even in ambiguous class distributions. Our experiments show that TabEBM generates synthetic data with higher quality and better statistical fidelity than existing methods. When used for data augmentation, our synthetic data consistently improves the classification performance across diverse datasets of various sizes, especially small ones. Code is available at https://github.com/andreimargeloiu/TabEBM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TabStruct: Measuring Structural Fidelity of Tabular DataXiangjian Jiang, Nikola Simidjievski, Mateja JamnikICLR 2026 · 被引用 10 次
- Huma: Censorship Circumvention via Web Protocol Tunneling with Deferred Traffic ReplacementSina Kamali, Diogo BarradasNDSS 2026 · 被引用 1 次
- Active Tabular Augmentation via Policy-Guided Diffusion InpaintingZheyu Zhang, Shuo Yang, Bardh Prenkaj, Gjergji KasneciICML 2026
- NRGBoost: Energy-Based Generative Boosted TreesJoão BravoICLR 2025
- Deep Incomplete Multi-View Clustering via Hierarchical Imputation and AlignmentYiming Du, Ziyu Wang, Jian Li, Rui Ning 等AAAI 2026
它引用的顶会 Paper20
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud 等ICLR 2020 · 被引用 643 次
- TabDDPM: Modelling Tabular Data with Diffusion ModelsAkim Kotelnikov, Dmitry Baranchuk, Ivan Rubachev, Artem BabenkoICML 2023 · 被引用 518 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Mixed-Type Tabular Data Synthesis with Score-based Diffusion in Latent SpaceHengrui Zhang, Jiani Zhang, Zhengyuan Shen, Balasubramaniam Srinivasan 等ICLR 2024 · 被引用 233 次
相关 Paper
- No MCMC for me: Amortized sampling for fast and stable training of energy-based modelsWill Sussman Grathwohl, Jacob Jin Kelly, Milad Hashemi, Mohammad Norouzi 等ICLR 2021 · 被引用 75 次
- Energy-Based Modelling for Discrete and Mixed Data via Heat Equations on Structured SpacesTobias Schröder, Zijing Ou, Yingzhen Li, Andrew B. DuncanNeurIPS 2024 · 被引用 5 次
- Synthetic Tabular Data Generation for Imbalanced Classification: The Surprising Effectiveness of an Overlap ClassAnnie D'souza, Swetha M, Sunita SarawagiAAAI 2025 · 被引用 9 次
- VAEBM: A Symbiosis between Variational Autoencoders and Energy-based ModelsZhisheng Xiao, Karsten Kreis, Jan Kautz, Arash VahdatICLR 2021 · 被引用 139 次
- EPIC: Effective Prompting for Imbalanced-Class Data Synthesis in Tabular Data Classification via Large Language ModelsJinhee Kim, Taesung Kim, Jaegul ChooNeurIPS 2024 · 被引用 22 次
