Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive Generation
Ying Li, Zefang Wang, Zhaode Wang, Zhiwen Chen, chengfei lv, Huan Wang
摘要
Scaling up visual autoregressive models improves generation quality but incurs substantial inference costs. Mixture-of-Experts (MoE) architectures mitigate this issue through sparse activation and have proven effective in large language models. However, training MoE models from scratch remains prohibitively expensive, and dense-to-MoE conversion for visual autoregressive models is still underexplored. To enable low-cost and high-quality dense-to-MoE conversion , we propose Prism-MoE , an efficient framework for transforming pretrained dense visual autoregressive models into sparse MoE models. Prism-MoE consists of two key components. First, we introduce trajectory-consistent Initialization, which formulates expert initialization as a principled decomposition problem and preserves the generation trajectory of pretrained models. Second, we propose a confidence-adaptive sparse fine-tuning framework that aligns expert specialization with the information density of visual tokens via confidence-aware routing supervision. Experiments show that Prism-MoE achieves dense-to-MoE conversion with less than 10% of the standard training budget, while maintaining generation quality comparable to dense baselines with only 37.5% active parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Autoregressive Image Generation using Residual QuantizationDoyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho 等CVPR 2022 · 被引用 184 次
相关 Paper
- Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts ConversionFilip Szatkowski, Bartosz Wójcik, Mikolaj Piórczynski, Simone ScardapaneNeurIPS 2024 · 被引用 19 次
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationYouwei Zheng, Yuxi Ren, Xin Xia, Xuefeng Xiao 等ICCV 2025 · 被引用 1 次
- Mixture of Tokens: Continuous MoE through Cross-Example AggregationSzymon Antoniak, Michal Krutul, Maciej Pióro, Jakub Krajewski 等NeurIPS 2024 · 被引用 6 次
- MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision TasksXingkui Zhu, Yiran Guan, Dingkang Liang, Yuchao Chen 等NeurIPS 2024 · 被引用 15 次
- DOT-MoE: Differentiable Optimal Transport for MoEficationUdbhav Bamba, Arnav Chavan, Aryamaan Thakur, Steven Teig 等ICML 2026
