Neodragon: Mobile Video Generation Using Diffusion Transformer
Animesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima, Adil Karjauv, Mohsen Ghafoorian, Amirhossein Habibian
Abstract
We introduce Neodragon, a text-to-video system capable of generating 2s (49 frames @24 fps) videos at a resolution of 640×1024 directly on a Qualcomm Hexagon NPU in a record ∼6.7s (7 FPS). Differing from existing transformer-based offline text-to-video generation models, Neodragon is the first to have been specifically optimised for mobile hardware to achieve efficient, low-cost, and high-fidelity video synthesis. We achieve this through four key technical contributions: (1) Replacing the original large 4.762B T5 XXL Text-Encoder with a much smaller 0.2B DT5 (DistilT5) with minimal quality loss, enabling the entire model to run without CPU offloading. This is enabled through a novel Text-Encoder Distillation procedure which uses only generative text-prompt data and does not require any image or video data. (2) Proposing an Asymmetric Decoder Distillation approach which allows us to replace the native codec-latent-VAE decoder with a more efficient one, without disturbing the generative latent-space of the video generation pipeline. (3) Pruning of MMDiT blocks within the denoiser backbone based on their relative importance, with recovery of original performance through a two-stage distillation process. (4) Reducing the NFE (Neural Functional Evaluation) requirement of the denoiser by performing step distillation using a technique adapted from DMD for pyramidal flow-matching, thereby significantly accelerating video generation. When paired with an optimised SSD1B first-frame image generator and QuickSRNet for 2× super-resolution, our end-to-end Neodragon system becomes a highly parameter (4.945B full model), memory (3.5GB peak RAM usage), and runtime (6.7s E2E latency) efficient mobile-friendly model, while achieving a VBench total score of 81.61, yielding high-fidelity generated videos. By enabling low-cost, private, and on-device text-to-video synthesis, Neodragon democratizes AI-based video content creation, empowering creators to generate high-quality videos without reliance on cloud services. Code and model will be made publicly available at our website.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29645bc9-b670-49a8-9062-3b6d65e09c8cCited by top-tier papers3
- ReHyAt: Recurrent Hybrid Attention for Video Diffusion TransformersMohsen Ghafoorian, Amirhossein HabibianCVPR 2026 · 5 citations
- PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient InferenceDenis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian et al.CVPR 2026 · 3 citations
- RFDM: Residual Flow Diffusion Models for Video EditingMohammadreza Salehi, Mehdi Noroozi, Luca Morreale, Ruchika Chavhan et al.CVPR 2026
Builds on33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- NanoFLUX: Distillation-Driven Compression of Large Text-to-Image Generation Models for Mobile DevicesRuchika Chavhan, Malcolm Chadwick, Alberto Gil Couto Pimentel Ramos, Luca Morreale et al.ICML 2026
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
- ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video GenerationZongyi Li, Shujie Hu, Shujie Liu, Long Zhou et al.ICLR 2025
- DOLLAR: Few-Step Video Generation Via Distillation and Latent Reward OptimizationZihan Ding, Chi Jin, Difan Liu, Haitian Zheng et al.ICCV 2025 · 3 citations
