VideoTetris: Towards Compositional Text-to-Video Generation
Ye Tian, Ling Yang, Haotian Yang, Yuan Gao, Yufan Deng, Xintao Wang, Zhaochen Yu, Xin Tao, Pengfei Wan, Di Zhang, Bin Cui
Abstract
Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose an enhanced video data preprocessing to enhance the training data regarding motion dynamics and prompt understanding, equipped with a new reference frame attention mechanism to improve the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71a8e466-1c7d-4e6b-9b3f-bc5d4533b806Cited by top-tier papers28
- StreamDiT: Real-Time Streaming Text-to-Video GenerationAkio Kodaira, Tingbo Hou, Ji Hou, Markos Georgopoulos et al.CVPR 2026 · 45 citations
- Token Perturbation Guidance for Diffusion ModelsJavad Rajabi, Soroush Mehraban, Seyedmorteza Sadat, Babak TaatiNeurIPS 2025 · 17 citations
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and GenerationLing Yang, Xinchen Zhang, Ye Tian, Shiyi Zhang et al.NeurIPS 2025 · 16 citations
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video GenerationWeinan Jia, Yuning Lu, Mengqi Huang, Hualiang Wang et al.ICLR 2026 · 14 citations
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu et al.ICLR 2026 · 13 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Grid Diffusion Models for Text-to-Video GenerationTaegyeong Lee, Soyeong Kwon, Taehwan KimCVPR 2024
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng et al.NeurIPS 2024 · 291 citations
- TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationXingrui Wang, Xin Li, Yaosi Hu, Hanxin Zhu et al.AAAI 2025 · 3 citations
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang et al.AAAI 2025 · 1 citation
- DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion PriorsJingyu Lin, Xinyi Shang, Peng Sun, Cunjian Chen et al.ICML 2026
