VIMI: Grounding Video Generation through Multi-modal Instruction
Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chieh Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov
摘要
Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual grounding and restricting their versatility and application in multimodal integration. To address this, we construct a large-scale multimodal prompt dataset by employing retrieval methods to pair incontext examples with the given text prompts and then utilize a two-stage training strategy to enable diverse video generation tasks within the same model. In the first stage, we propose a multimodal conditional video generation framework for pretraining on these augmented datasets, establishing a foundational model for grounded video generation. Secondly, we finetune the model from the first stage on three video generation tasks, incorporating multimodal instructions. This process further refines the model's ability to handle diverse inputs and tasks, ensuring seamless integration of multimodal information. After this two-stage training process, VIMI demonstrates multimodal understanding capabilities, producing contextually rich and personalized videos grounded in the provided inputs, as shown in Figure 1 . Compared to previous visual grounded video generation methods, VIMI can synthesize consistent and temporally coherent videos with large motion while retaining the semantic control. Lastly, VIMI also achieves state-of-theart text-to-video generation results on UCF101 benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control ConditionsYuanhao Cai, He Zhang, Xi Chen, Jinbo Xing 等NeurIPS 2025 · 被引用 19 次
- AlcheMinT: Fine-grained Temporal Control for Multi-Reference Consistent Video GenerationSharath Girish, Viacheslav Ivanov, Tsai-Shien Chen, Hao Chen 等CVPR 2026 · 被引用 4 次
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept PersonalizationTsai-Shien Chen, Aliaksandr Siarohin, Gordon Guocheng Qian, Kuan-Chieh Jackson Wang 等CVPR 2026 · 被引用 4 次
- Multi-subject Open-set Personalization in Video GenerationTsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang 等CVPR 2025
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video GenerationWenhao Wang, Yi YangICCV 2025 · 被引用 1 次
- UNIMO-G: Unified Image Generation through Multimodal Conditional DiffusionWei Li, Xue Xu, Jiachen Liu, Xinyan XiaoACL 2024 · 被引用 5 次
- JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and GenerationKai Liu, Jungang Li, Yuchong Sun, Shengqiong Wu 等NeurIPS 2025 · 被引用 18 次
- V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction TuningHang Hua, Yunlong Tang, Chenliang Xu, Jiebo LuoAAAI 2025 · 被引用 61 次
