Aid: Adapting Image2video Diffusion Models for Instruction-Guided Video Prediction
Zhen Xing, Qi Dai, Zejia Weng, Zuxuan Wu, Yu-Gang Jiang
Abstract
Text-guided video prediction (TVP) involves predicting the motion of future frames from the initial frame according to an instruction, which has wide applications in virtual reality, robotics, and content creation. Previous TVP methods make significant breakthroughs by adapting Stable Diffusion for this task. However, they struggle with frame consistency and temporal stability primarily due to the limited scale of video datasets. We observe that pretrained Image2Video diffusion models possess good priors for video dynamics but they lack textual control. Hence, transferring Image2Video models to leverage their video dynamic priors while injecting instruction control to generate controllable videos is both a meaningful and challenging task. To achieve this, we introduce the Multi-Modal Large Language Model (MLLM) to predict future video states based on initial frames and text instructions. More specifically, we design a dual query transformer (DQFormer) architecture, which integrates the instructions and frames into the conditional embeddings for future frame prediction. Additionally, we develop Long-Short Term Temporal Adapters and Spatial Adapters that can quickly transfer general video diffusion models to specific scenarios with minimal training costs. Experimental results show that our method significantly outperforms state-of-the-art techniques on four datasets: Something Something V2, Epic Kitchen-100, Bridge Data, and UCF-101. Notably, AID achieves 91.2% and 55.5% FVD improvements on Bridge and SSv2 respectively, demonstrating its effectiveness in various domains. More examples can be found at our website https://chenhsing.github.io/AID.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feb2c1f7-eb41-45ce-9369-5d87eb8a5f92Cited by top-tier papers7
- Human2Robot: Learning Robot Actions from Paired Human-Robot VideosSicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing et al.AAAI 2026 · 15 citations
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang et al.ICCV 2025 · 10 citations
- CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image GenerationHui Zhang, Dexiang Hong, Yitong Wang, Jie Shao et al.ICCV 2025 · 7 citations
- DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion AdaptationZun Wang, Jialu Li, Han Lin, Jaehong Yoon et al.AAAI 2026 · 5 citations
- FlashMotion: Few-Step Controllable Video Generation with Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Haidong Cao et al.CVPR 2026 · 5 citations
Builds on45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Seer: Language Instructed Video Prediction with Latent Diffusion ModelsXianfan Gu, Chuan Wen, Weirui Ye, Jiaming Song et al.ICLR 2024 · 57 citations
- Modular-Cam: Modular Dynamic Camera-view Video Generation with LLMZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen et al.AAAI 2025 · 6 citations
- Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video GenerationTsu-Jui Fu, Licheng Yu, Ning Zhang, Cheng-Yang Fu et al.CVPR 2023
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li et al.CVPR 2026 · 5 citations
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar et al.AAAI 2025 · 3 citations
