A Learnable Discrete-Prior Fusion Autoencoder with Contrastive Learning for Tabular Data Synthesis
Rongchao Zhang, Yiwei Lou, Dexuan Xu, Yongzhi Cao, Hanpin Wang, Yu Huang
Abstract
The actual collection of tabular data for sharing involves confidentiality and privacy constraints, leaving the potential risks of machine learning for interventional data analysis unsafely averted. Synthetic data has emerged recently as a privacyprotecting solution to address this challenge. However, existing approaches regard discrete and continuous modal features as separate entities, thus falling short in properly capturing their inherent correlations. In this paper, we propose a novel contrastive learning guided Gaussian Transformer autoencoder, termed GTCoder, to synthesize photo-realistic multimodal tabular data for scientific research. Our approach introduces a transformer-based fusion module that seamlessly integrates multimodal features, permitting for mining more informative latent representations. The attention within the fusion module directs the integrated output features to focus on critical components that facilitate the task of generating latent embeddings. Moreover, we formulate a contrastive learning strategy to implicitly constrain the embeddings from discrete features in the latent feature space by encouraging the similar discrete feature distributions closer while pushing the dissimilar further away, in order to better enhance the representation of the latent embedding. Experimental results indicate that GTCoder is effective to generate photo-realistic synthetic data, with interactive interpretation of latent embedding, and performs favorably against some baselines on most real-world and simulated datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53a55d5f-0dd4-42f2-9cdc-2b24a8629d04Cited by top-tier papers3
- MoleBridge: Synthetic Space Projecting with Discrete Markov BridgesRongchao Zhang, Yu Huang, Yongzhi Cao, Hanpin WangNeurIPS 2025 · 8 citations
- Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar DiagnosisChen Feng, Zhuo Zhi, Zhao Huang, Jiawei Ge et al.CVPR 2026 · 4 citations
- Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion BridgeRongchao Zhang, Chengxin Li, Yiwei Lou, Yuling Shi et al.CVPR 2026 · 1 citation
Builds on18
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- HyperTransformer: A Textural and Spectral Feature Fusion Transformer for PansharpeningWele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2022 · 175 citations
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang et al.EMNLP 2022 · 148 citations
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist et al.ICLR 2021 · 86 citations
Related papers
- CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular SynthesisChaejeong Lee, Jayoung Kim, Noseong ParkICML 2023 · 97 citations
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- Multimodal Adversarially Learned Inference with Factorized DiscriminatorsWenxue Chen, Jianke ZhuAAAI 2022 · 3 citations
- CG-TGAN: Conditional Generative Adversarial Networks with Graph Neural Networks for Tabular Data SynthesizingSeungcheol Lee, Moohong MinAAAI 2025 · 4 citations
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang et al.AAAI 2026
