FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention
Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Qiang Xu
Abstract
A woman carrying a bouquet of vibrant flowers walks along the beach A man and a woman are standing together in a forested area
A young woman wearing a white t-shirt is sitting indoors.
Figure 1. FullDiT is a multi-task video generative foundation model that unifies conditional learning with full self-attention. With self-attention's long-context learning ability, FullDiT can flexibly take different combinations of input to generate high-quality videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac6c822b-3415-4124-bfe9-80e55c7d2874Cited by top-tier papers3
- OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video GenerationDonghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li et al.ICML 2026 · 3 citations
- EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect CreationShiyuan Yang, Ruihuang Li, Jiale Tao, Shuai Shao et al.CVPR 2026 · 2 citations
- DiasR: Dual-Modal Identity-Anchored Sparse Routing for Efficient Multi-Subject Video GenerationYang-yang Li, Wu Liu, Jie Li, Xinchen Liu et al.ICML 2026
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video GenerationMinghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu et al.CVPR 2025
- UniVideo: Unified Understanding, Generation, and Editing for VideosCong Wei, Quande Liu, Zixuan Ye, Qiulin Wang et al.ICLR 2026 · 90 citations
- FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less ComputeSotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler et al.CVPR 2025
- LayerT2V: A Unified Multi-Layer Video Generation FrameworkGuangzhao Li, Kangrui Cen, Baixuan Zhao, Yi Xin et al.ICML 2026 · 2 citations
- Unified In-Context Video EditingZixuan Ye, Xuanhua He, Quande Liu, Qiulin Wang et al.ICLR 2026 · 37 citations
