A Theory for Conditional Generative Modeling on Multiple Data Sources
Rongzhen Wang, Yan Zhang, Chenyu Zheng, Chongxuan Li, Guoqiang Wu
Abstract
The success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities. However, the interaction among these sources remains theoretically underexplored. This paper takes the first step toward a rigorous analysis of multi-source training in conditional generative modeling, where each condition represents a distinct data source. Specifically, we establish a general distribution estimation error bound in average total variation distance for conditional maximum likelihood estimation based on the bracketing number. Our result shows that when source distributions share certain similarities and the model is expressive enough, multi-source training guarantees a sharper bound than single-source training. We further instantiate the general theory on conditional Gaussian estimation and deep generative models including autoregressive and flexible energybased models, by characterizing their bracketing numbers. The results highlight that the number of sources and similarity among source distributions improve the advantage of multi-source training. Simulations and real-world experiments are conducted to validate the theory, with code available at: https://github.com/ML-GSAI/ Multi-Source-GM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Scaling Diffusion Transformers Efficiently via μPChenyu Zheng, Xinyu Zhang, Rongzhen Wang, Wei Huang et al.NeurIPS 2025 · 7 citations
- Metadata Conditioning Accelerates Language Model Pre-trainingTianyu Gao, Alexander Wettig, Luxi He, Yihe Dong et al.ICML 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
Related papers
- A Discriminative Technique for Multiple-Source AdaptationCorinna Cortes, Mehryar Mohri, Ananda Theertha Suresh, Ningshan ZhangICML 2021 · 15 citations
- When Does Closeness in Distribution Imply Representational Similarity? An Identifiability PerspectiveBeatrix M. G. Nielsen, Emanuele Marconato, Andrea Dittadi, Luigi GreseleNeurIPS 2025 · 7 citations
- Learning a 1-layer conditional generative model in total variationAjil Jalal, Justin Singh Kang, Ananya Uppal, Kannan Ramchandran et al.NeurIPS 2023
- Learning Joint Latent Space EBM Prior Model for Multi-layer GeneratorJiali Cui, Ying Nian Wu, Tian HanCVPR 2023
- A Likelihood Based Approach to Distribution Regression Using Conditional Deep Generative ModelsShivam Kumar, Yun Yang, Lizhen LinICML 2025
