FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
Sotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Albert Pumarola, Ali K. Thabet, Edgar Schönfeld
2025Year
10Top-tier citations
Abstract
agnostic to input and conditioning modalities. We show how our approach can be readily extended for video generation, where FlexiDiT models generate samples with up to 75% less compute without compromising performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97ca5a87-c4fb-4944-8260-754c789f7fdfCited by top-tier papers10
- Scale-wise Distillation of Diffusion ModelsNikita Starodubcev, Ilya Drobyshevskiy, Denis Kuznedelev, Artem Babenko et al.ICLR 2026 · 13 citations
- SeaCache: Spectral-Evolution-Aware Cache for Accelerating Diffusion ModelsJiwoo Chung, Sangeek Hyun, MinKyu Lee, Byeongju Han et al.CVPR 2026 · 9 citations
- Foresight: Adaptive Layer Reuse for Accelerated and High-Quality Text-to-Video GenerationMuhammad Adnan, Nithesh Kurella, Akhil Arunkumar, Prashant J. NairNeurIPS 2025 · 6 citations
- RAPID: Tri-Level Reinforced Acceleration Policies for Diffusion TransformerWangbo Zhao, Yizeng Han, Zhiwei Tang, Jiasheng Tang et al.ICLR 2026 · 5 citations
- InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster GenerationYuxin Qin, Ke Cao, Haowei Liu, Ao Ma et al.CVPR 2026 · 5 citations
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video LearningA. J. Piergiovanni, Weicheng Kuo, Anelia AngelovaCVPR 2023
- Dream Video: Composing Your Dream Videos with Customized Subject and MotionYujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan et al.CVPR 2024
- MAGVIT: Masked Generative Video TransformerLijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama et al.CVPR 2023
- One Algorithm to Align Them AllBoyi Pang, Savva Ignatyev, Vladimir Ippolitov, Ramil Khafizov et al.CVPR 2026
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin et al.ICCV 2025
