Mind the Time: Temporally-Controlled Multi-Event Video Generation
Ziyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Yuwei Fang, Varnith Chordia, Igor Gilitschenski, Sergey Tulyakov
Abstract
Prompts: "A cat is on a table" → "jumps to the floor" → "jumps to the sofa" → "walks to the table again" → "sits down and looks around" Prompts: "A man is typing on a laptop" → "touches his headphone with his right hand" → "closes the laptop with his left hand" → "stands up" Prompts: "A man is smiling" → "looks to his left with a surprised face" → "lowers his head with a sad face" → "smiles to the camera again" Prompts: "An old lady waves her right hand" → "makes a thumbs-up gesture" → "makes a heart gesture" → "gives a blow kiss" Figure 1 . Time-controlled multi-event video generation with MinT. Given a sequence of event text prompts and their desired start and end timestamps, MinT synthesizes smoothly connected events with consistent subjects and backgrounds. In addition, it can control the time span of each event flexibly. Here, we show the results of sequential gestures, daily activities, facial expressions, and cat movements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9b790d8-62f6-4a42-8f51-0dab38dbdd3aCited by top-tier papers17
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein et al.NeurIPS 2025 · 132 citations
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace et al.NeurIPS 2025 · 36 citations
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou et al.CVPR 2026 · 33 citations
- EchoShot: Multi-Shot Portrait Video GenerationJiahao Wang, Hualian Sheng, Sijia Cai, Weizhan Zhang et al.NeurIPS 2025 · 30 citations
- EasyV2V: A High-quality Instruction-based Video Editing FrameworkJinjie Mai, Chaoyang Wang, Gordon Guocheng Qian, Willi Menapace et al.CVPR 2026 · 12 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video GenerationMinghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu et al.CVPR 2025
- SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile DeviceYushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu et al.CVPR 2025
- MoCha: Towards Movie-Grade Talking Character GenerationCong Wei, Bo Sun, Haoyu Ma, Ji Hou et al.NeurIPS 2025 · 2 citations
- MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose GuidanceYuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang et al.ICML 2025
- VideoBooth: Diffusion-based Video Generation with Image PromptsYuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si et al.CVPR 2024
