Mind the Time: Temporally-Controlled Multi-Event Video Generation
Ziyi Wu, Aliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Yuwei Fang, Varnith Chordia, Igor Gilitschenski, Sergey Tulyakov
摘要
Prompts: "A cat is on a table" → "jumps to the floor" → "jumps to the sofa" → "walks to the table again" → "sits down and looks around" Prompts: "A man is typing on a laptop" → "touches his headphone with his right hand" → "closes the laptop with his left hand" → "stands up" Prompts: "A man is smiling" → "looks to his left with a surprised face" → "lowers his head with a sad face" → "smiles to the camera again" Prompts: "An old lady waves her right hand" → "makes a thumbs-up gesture" → "makes a heart gesture" → "gives a blow kiss" Figure 1 . Time-controlled multi-event video generation with MinT. Given a sequence of event text prompts and their desired start and end timestamps, MinT synthesizes smoothly connected events with consistent subjects and backgrounds. In addition, it can control the time span of each event flexibly. Here, we show the results of sequential gestures, daily activities, facial expressions, and cat movements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein 等NeurIPS 2025 · 被引用 132 次
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion ModelsZiyi Wu, Anil Kag, Ivan Skorokhodov, Willi Menapace 等NeurIPS 2025 · 被引用 36 次
- OneStory: Coherent Multi-Shot Video Generation with Adaptive MemoryZhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 等CVPR 2026 · 被引用 33 次
- EchoShot: Multi-Shot Portrait Video GenerationJiahao Wang, Hualian Sheng, Sijia Cai, Weizhan Zhang 等NeurIPS 2025 · 被引用 30 次
- EasyV2V: A High-quality Instruction-based Video Editing FrameworkJinjie Mai, Chaoyang Wang, Gordon Guocheng Qian, Willi Menapace 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video GenerationMinghong Cai, Xiaodong Cun, Xiaoyu Li, Wenze Liu 等CVPR 2025
- SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile DeviceYushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu 等CVPR 2025
- MoCha: Towards Movie-Grade Talking Character GenerationCong Wei, Bo Sun, Haoyu Ma, Ji Hou 等NeurIPS 2025 · 被引用 2 次
- MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose GuidanceYuang Zhang, Jiaxi Gu, Li-Wen Wang, Han Wang 等ICML 2025
- VideoBooth: Diffusion-based Video Generation with Image PromptsYuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si 等CVPR 2024
