Mask^2DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation
Tianhao Qi, Jianlong Yuan, Wanquan Feng, Shancheng Fang, Jiawei Liu, SiYu Zhou, Qian He, Hongtao Xie, Yongdong Zhang
摘要
Sora has unveiled the immense potential of the Diffusion Transformer (DiT) architecture in single-scene video generation. However, the more challenging task of multi-scene video generation, which offers broader applications, remains relatively underexplored. To bridge this gap, we propose Mask 2 DiT, a novel approach that establishes finegrained, one-to-one alignment between video segments and their corresponding text annotations. Specifically, we introduce a symmetric binary mask at each attention layer within the DiT architecture, ensuring that each text annotation applies exclusively to its respective video segment while preserving temporal coherence across visual tokens. This attention mechanism enables precise segment-level textualto-visual alignment, allowing the DiT architecture to effectively handle video generation tasks with a fixed number of scenes. To further equip the DiT architecture with the ability to generate additional scenes based on existing ones, we incorporate a segment-level conditional mask, which conditions each newly generated segment on the preceding video segments, thereby enabling auto-regressive scene extension. Both qualitative and quantitative experiments confirm that Mask 2 DiT excels in maintaining visual consistency across segments while ensuring semantic alignment between each segment and its corresponding text description. Our project page is https://tianhao-qi.github.io/Mask2DiTProject/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MultiShotMaster: A Controllable Multi-Shot Video Generation FrameworkQinghe Wang, Xiaoyu Shi, Baolu Li, Weikang Bian 等CVPR 2026 · 被引用 33 次
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
- CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion ModelsXiaoxue Wu, Bingjie Gao, Yu Qiao, Yaohui Wang 等ICLR 2026 · 被引用 26 次
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu 等ICLR 2026 · 被引用 13 次
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic TransitionsXiaoxue Wu, Xinyuan Chen, Yaohui Wang, Yu QiaoCVPR 2026 · 被引用 5 次
它引用的顶会 Paper21
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 等ICML 2024 · 被引用 464 次
相关 Paper
- LayerT2V: A Unified Multi-Layer Video Generation FrameworkGuangzhao Li, Kangrui Cen, Baixuan Zhao, Yi Xin 等ICML 2026 · 被引用 2 次
- VDT: General-purpose Video Diffusion Transformers via Mask ModelingHaoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo 等ICLR 2024 · 被引用 117 次
- AV-DiT: Taming Image Diffusion Transformers for Efficient Joint Audio and Video GenerationKai Wang, Shijian Deng, Jing Shi, Dimitrios Hatzinakos 等ACM MM 2025 · 被引用 2 次
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon 等NeurIPS 2025 · 被引用 6 次
- Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial GenerationYushe Cao, Dianxi Shi, Xing Fu, Xuechao Zou 等AAAI 2026
