Lune

CVPR2025Top-tier venue

DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Minghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu, Zhaoyang Zhang, Yong Zhang, Ying Shan, Xiangyu Yue

2025Year
34Top-tier citations

Abstract

SHIAE, CUHK (b) "Frosty pine: close-up shot medium shot forest vista, cinematic" (a) "Athlete glides across ocean waters snow mountain sand dunes" Figure 1. Our method DiTCtrl takes multiple text prompts as input and demonstrates superior capability in generating longer videos with multiple events, long-range coherence and smooth transitions as output.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9ce5bea8-e34f-4907-8796-8cd81e6b3e91

Cited by top-tier papers34

Ask how each one uses it

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines