VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data
Chao Ma, Sebastian Tschiatschek, Richard E. Turner, José Miguel Hernández-Lobato, Cheng Zhang
Abstract
Deep generative models often perform poorly in real-world applications due to the heterogeneity of natural data sets. Heterogeneity arises from data containing different types of features (categorical, ordinal, continuous, etc.) and features of the same type having different marginal distributions. We propose an extension of variational autoencoders (VAEs) called VAEM to handle such heterogeneous data. VAEM is a deep generative model that is trained in a two stage manner such that the first stage provides a more uniform representation of the data to the second stage, thereby sidestepping the problems caused by heterogeneous data. We provide extensions of VAEM to handle partially observed data, and demonstrate its performance in data generation, missing data prediction and sequential feature selection tasks. Our results show that VAEM broadens the range of real-world applications where deep generative models can be successfully deployed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9760421a-00a3-43d6-9ddc-33d7075d8c3cCited by top-tier papers16
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- Learning to Maximize Mutual Information for Dynamic Feature SelectionIan Connick Covert, Wei Qiu, Mingyu Lu, Nayoon Kim et al.ICML 2023 · 67 citations
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 56 citations
- Mitigating Modality Collapse in Multimodal VAEs via Impartial OptimizationAdrián Javaloy, Maryam Meghdadi, Isabel ValeraICML 2022 · 49 citations
- Language Models are Realistic Tabular Data GeneratorsVadim Borisov, Kathrin Seßler, Tobias Leemann, Martin Pawelczyk et al.ICLR 2023 · 45 citations
Related papers
- Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte CarloIgnacio Peis, Chao Ma, José Miguel Hernández-LobatoNeurIPS 2022 · 25 citations
- Variational Inference for Discriminative Learning with Generative Modeling of Feature IncompletionKohei Miyaguchi, Takayuki Katsuki, Akira Koseki, Toshiya IwamoriICLR 2022 · 4 citations
- A Critical Look at the Consistency of Causal Estimation with Deep Latent Variable ModelsSeveri Rissanen, Pekka MarttinenNeurIPS 2021 · 38 citations
- BooVAE: Boosting Approach for Continual Learning of VAEEvgenii Egorov, Anna Kuzina, Evgeny BurnaevNeurIPS 2021 · 34 citations
- Unbiased learning of deep generative models with structured discrete representationsHenry C. Bendekgey, Gabe Hope, Erik B. SudderthNeurIPS 2023 · 2 citations
