Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT
Le Zhuo, Ruoyi Du, Han Xiao, Yangguang Li, Dongyang Liu, Rongjie Huang, Wenze Liu, Xiangyang Zhu, Fu-Yun Wang, Zhanyu Ma, Xu Luo, Zehan Wang
摘要
Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduced a sigmoid time discretization schedule to reduce sampling steps in solving the Flow ODE and the Context Drop method to merge redundant visual tokens for faster network evaluation, effectively boosting the overall sampling speed. Thanks to these improvements, Lumina-Next not only improves the quality and efficiency of basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities and multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-view, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights, we aim to advance the development of next-generation generative AI capable of universal modeling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- OmniGen2: Towards Instruction-Aligned Multimodal GenerationChenyuan Wu, Jiahao Wang, Pengfei Zheng, Ruiran Yan 等CVPR 2026 · 被引用 231 次
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang 等ICLR 2026 · 被引用 181 次
- DDT: Decoupled Diffusion TransformerShuai Wang, Zhi Tian, Weilin Huang, Limin WangCVPR 2026 · 被引用 102 次
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen 等NeurIPS 2024 · 被引用 86 次
- PixelDiT: Pixel Diffusion Transformers for Image GenerationYongsheng Yu, Wei Xiong, Weili Nie, Yichen Sheng 等CVPR 2026 · 被引用 82 次
它引用的顶会 Paper40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Lumina-T2X: Scalable Flow-based Large Diffusion Transformer for Flexible Resolution GenerationPeng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du 等ICLR 2025
- Lumina-Image 2.0: a Unified and Efficient Image Generative FrameworkQi Qin, Le Zhuo, Yi Xin, Ruoyi Du 等ICCV 2025 · 被引用 11 次
- dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language ModelsYi Xin, Siqi Luo, Tianxiang Xu, Qi Qin 等CVPR 2026 · 被引用 3 次
- Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model PerspectiveHangjie Yuan, Weihua Chen, Jun Cen, Hu Yu 等ICLR 2026 · 被引用 21 次
- LaVin-DiT: Large Vision Diffusion TransformerZhaoqing Wang, Xiaobo Xia, Runnan Chen, Dongdong Yu 等CVPR 2025
