DiMSUM: Diffusion Mamba - A Scalable and Unified Spatial-Frequency Method for Image Generation
Hao Phung, Quan Dao, Trung Tuan Dao, Viet Hoang Phan, Dimitris N. Metaxas, Anh Tuan Tran
摘要
We introduce a novel state-space architecture for diffusion models, effectively harnessing spatial and frequency information to enhance the inductive bias towards local features in input images for image generation tasks. While state-space networks, including Mamba, a revolutionary advancement in recurrent neural networks, typically scan input sequences from left to right, they face difficulties in designing effective scanning strategies, especially in the processing of image data. Our method demonstrates that integrating wavelet transformation into Mamba enhances the local structure awareness of visual inputs and better captures long-range relations of frequencies by disentangling them into wavelet subbands, representing both low- and high-frequency components. These wavelet-based outputs are then processed and seamlessly fused with the original Mamba outputs through a cross-attention fusion layer, combining both spatial and frequency information to optimize the order awareness of state-space models which is essential for the details and overall quality of image generation. Besides, we introduce a globally-shared transformer to supercharge the performance of Mamba, harnessing its exceptional power to capture global relationships. Through extensive experiments on standard benchmarks, our method demonstrates superior results compared to DiT and DIFFUSSM, achieving faster training convergence and delivering high-quality outputs. The codes and pretrained models are released at https://github.com/VinAIResearch/DiMSUM.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LiT: Delving into a Simple Linear Diffusion Transformer for Image GenerationJiahao Wang, Ning Kang, Lewei Yao, Mengzhao Chen 等ICCV 2025 · 被引用 10 次
- DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion ModelingYuang Ai, Qihang Fan, Xuefeng Hu, Zhenheng Yang 等NeurIPS 2025 · 被引用 8 次
- RiverMamba: A State Space Model for Global River Discharge and Flood ForecastingMohamad Hakam Shams Eddin, Yikui Zhang, Stefan Kollet, Jürgen GallNeurIPS 2025 · 被引用 7 次
- Stochastic Layer-Wise Shuffle for Improving Vision Mamba TrainingZizheng Huang, Haoxing Chen, Jiaqi Li, Jun Lan 等ICML 2025
- Improved Training Technique for Latent Consistency ModelsQuan Dao, Khanh Doan, Di Liu, Trung Le 等ICLR 2025
它引用的顶会 Paper38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- DiM-TS: Bridge the Gap Between Selective State Space Models and Time Series for Generative ModelingZihao Yao, Jiankai Zuo, Yaying ZhangAAAI 2026
- MobileMamba: Lightweight Multi-Receptive Visual Mamba NetworkHaoyang He, Jiangning Zhang, Yuxuan Cai, Hongxu Chen 等CVPR 2025
- A Novel State Space Model with Local Enhancement and State Sharing for Image FusionZihan Cao, Xiao Wu, Liang-Jian Deng, Yu ZhongACM MM 2024 · 被引用 24 次
- Mamba Modulation: On the Length Generalization of Mamba ModelsPeng Lu, Jerry Huang, Qiuhao Zeng, Xinyu Wang 等NeurIPS 2025 · 被引用 2 次
- SaMam: Style-aware State Space Model for Arbitrary Image Style TransferHongda Liu, Longguang Wang, Ye Zhang, Ziru Yu 等CVPR 2025
