MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion
Onkar Kishor Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal, Rekha Singhal
2025Year
2Top-tier citations
Abstract
We introduce MotionAura, a novel Text-to-Video generation model that predicts discrete tokens obtained from our large scale pre-trained 3D VAE. The displayed frames represent videos generated by our model when provided with the captions shown below each frame. The following link hosts the above generated videos along with other samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca924e17-96bc-4d67-8064-e0a56d7338d7Cited by top-tier papers2
- PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and GenerationOnkar Susladkar, Tushar Prakash, Adheesh Sunil Juvekar, Kiet A. Nguyen et al.CVPR 2026 · 6 citations
- Di[M]O: Distilling Masked Diffusion Models Into One-Step GeneratorYuanzhi Zhu, Xi Wang, Stéphane Lathuilière, Vicky KalogeitonICCV 2025
Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Make It Move: Controllable Image-to-Video Generation with Text DescriptionsYaosi Hu, Chong Luo, Zhenzhong ChenCVPR 2022 · 56 citations
- CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerZhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding et al.ICLR 2025
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAEYazhou Xing, Yang Fei, Yingqing He, Jingye Chen et al.ICCV 2025 · 2 citations
- Towards Robust and Controllable Text-to-Motion via Masked Autoregressive DiffusionZongye Zhang, Bohan Kong, Qingjie Liu, Yunhong WangACM MM 2025 · 2 citations
