A Learnable Discrete-Prior Fusion Autoencoder with Contrastive Learning for Tabular Data Synthesis
Rongchao Zhang, Yiwei Lou, Dexuan Xu, Yongzhi Cao, Hanpin Wang, Yu Huang
摘要
The actual collection of tabular data for sharing involves confidentiality and privacy constraints, leaving the potential risks of machine learning for interventional data analysis unsafely averted. Synthetic data has emerged recently as a privacyprotecting solution to address this challenge. However, existing approaches regard discrete and continuous modal features as separate entities, thus falling short in properly capturing their inherent correlations. In this paper, we propose a novel contrastive learning guided Gaussian Transformer autoencoder, termed GTCoder, to synthesize photo-realistic multimodal tabular data for scientific research. Our approach introduces a transformer-based fusion module that seamlessly integrates multimodal features, permitting for mining more informative latent representations. The attention within the fusion module directs the integrated output features to focus on critical components that facilitate the task of generating latent embeddings. Moreover, we formulate a contrastive learning strategy to implicitly constrain the embeddings from discrete features in the latent feature space by encouraging the similar discrete feature distributions closer while pushing the dissimilar further away, in order to better enhance the representation of the latent embedding. Experimental results indicate that GTCoder is effective to generate photo-realistic synthetic data, with interactive interpretation of latent embedding, and performs favorably against some baselines on most real-world and simulated datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MoleBridge: Synthetic Space Projecting with Discrete Markov BridgesRongchao Zhang, Yu Huang, Yongzhi Cao, Hanpin WangNeurIPS 2025 · 被引用 8 次
- Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar DiagnosisChen Feng, Zhuo Zhi, Zhao Huang, Jiawei Ge 等CVPR 2026 · 被引用 4 次
- Steering Where to Diffuse: Generative Modeling of Phenotypic Response Simulation with Steered Diffusion BridgeRongchao Zhang, Chengxin Li, Yiwei Lou, Yuling Shi 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper18
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- HyperTransformer: A Textural and Spectral Feature Fusion Transformer for PansharpeningWele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2022 · 被引用 175 次
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang 等EMNLP 2022 · 被引用 148 次
- Semantic Re-tuning with Contrastive TensionFredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist 等ICLR 2021 · 被引用 86 次
相关 Paper
- CoDi: Co-evolving Contrastive Diffusion Models for Mixed-type Tabular SynthesisChaejeong Lee, Jayoung Kim, Noseong ParkICML 2023 · 被引用 97 次
- CTSyn: A Foundation Model for Cross Tabular Data GenerationXiaofeng Lin, Chenheng Xu, Matthew Yang, Guang ChengICLR 2025
- Multimodal Adversarially Learned Inference with Factorized DiscriminatorsWenxue Chen, Jianke ZhuAAAI 2022 · 被引用 3 次
- CG-TGAN: Conditional Generative Adversarial Networks with Graph Neural Networks for Tabular Data SynthesizingSeungcheol Lee, Moohong MinAAAI 2025 · 被引用 4 次
- Multimodal Gaussian Mixture Variational Autoencoder with Consistency RegularizationsYarui Chen, Lehan Hong, Jianlin Shao, Jianning Yang 等AAAI 2026
