Multi-Subspace Multi-Modal Modeling for Diffusion Models: Estimation, Convergence and Mixture of Experts
Ruofeng Yang, Yongcan Li, Bo Jiang, Cheng Chen, Shuai Li
Abstract
Recently, diffusion models have achieved a great performance with a small dataset of size and a fast optimization process. However, the estimation error of diffusion models suffers from the curse of dimensionality with the data dimension . Since images are usually a union of low-dimensional manifolds, current works model the data as a union of linear subspaces with Gaussian latent and achieve a bound. Though this modeling reflects the multi-manifold property, the Gaussian latent can not capture the multi-modal property of the latent manifold. To bridge this gap, we propose the mixture subspace of low-rank mixture of Gaussian (MoLR-MoG) modeling, which models the target data as a union of linear subspaces, and each subspace admits a mixture of Gaussian latent ( modals with dimension ). With this modeling, the corresponding score function naturally has a mixture of expert (MoE) structure, captures the multi-modal information, and contains nonlinear property. We first conduct real-world experiments to show that the generation results of MoE-latent MoG NN are much better than MoE-latent Gaussian score. Furthermore, MoE-latent MoG NN achieves a comparable performance with MoE-latent Unet with parameters. These results indicate that the MoLR-MoG modeling is reasonable and suitable for real-world data. After that, based on such MoE-latent MoG score, we provide a estimation error, which escapes the curse of dimensionality by using data structure. Finally, we study the optimization process and prove the convergence guarantee under the MoLR-MoG modeling. Combined with these results, under a setting close to real-world data, this work explains why diffusion models only require a small training sample and enjoy a fast optimization process to achieve a great performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef94f756-d664-4c44-b073-16004a099a8eCited by top-tier papers2
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and MethodHaochen Wang, Xiangtai Li, Zilong Huang, Anran Wang et al.ICLR 2026 · 45 citations
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsHaochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou et al.ICLR 2026 · 18 citations
Builds on23
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Score Approximation, Estimation and Distribution Recovery of Diffusion Models on Low-Dimensional DataMinshuo Chen, Kaixuan Huang, Tuo Zhao, Mengdi WangICML 2023 · 168 citations
Related papers
- On the Expressive Power of Mixture-of-Experts for Structured Complex TasksMingze Wang, Weinan ENeurIPS 2025 · 3 citations
- Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal DistributionsJingda Wu, Changxiao CaiICML 2026 · 1 citation
- MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic GuidanceHang Xu, Yang Xiao, Changlong Jiang, Haohong Kuang et al.AAAI 2026
- Mixture-of-Subspaces in Low-Rank AdaptationTaiqiang Wu, Jiahao Wang, Zhe Zhao, Ngai WongEMNLP 2024 · 14 citations
- Data Representations' Study of Latent Image ManifoldsIlya Kaufman, Omri AzencotICML 2023 · 11 citations
