Lune

ICCV2025Top-tier venue

FullDiT: Video Generative Foundation Models with Multimodal Control via Full Attention

Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Qiang Xu

2025Year
5Citations
3Top-tier citations

Abstract

A woman carrying a bouquet of vibrant flowers walks along the beach A man and a woman are standing together in a forested area

A young woman wearing a white t-shirt is sitting indoors.

Figure 1. FullDiT is a multi-task video generative foundation model that unifies conditional learning with full self-attention. With self-attention's long-context learning ability, FullDiT can flexibly take different combinations of input to generate high-quality videos.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ac6c822b-3415-4124-bfe9-80e55c7d2874

Cited by top-tier papers3

Ask how each one uses it

Builds on40

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines