SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
Koichi Namekata, Sherwin Bahmani, Ziyi Wu, Yash Kant, Igor Gilitschenski, David B. Lindell
摘要
Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error, e.g., involving re-generating videos with different random seeds. Recent techniques address this issue by fine-tuning a pre-trained model to follow conditioning signals, such as bounding boxes or point trajectories. Yet, this fine-tuning procedure can be computationally expensive, and it requires datasets with annotated object motion, which can be difficult to procure. In this work, we introduce SG-I2V, a framework for controllable image-to-video generation that is self-guided—offering zero-shot control by relying solely on the knowledge present in a pre-trained image-to-video diffusion model without the need for fine-tuning or external knowledge. Our zero-shot method outperforms unsupervised baselines while significantly narrowing down the performance gap with supervised models in terms of visual quality and motion fidelity.Additional details and video results are available on our project page: https://kmcode1.github.io/Projects/SG-I2V
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control SignalsNate Gillman, Charles Herrmann, Michael Freeman, Daksh Aggarwal 等NeurIPS 2025 · 被引用 61 次
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRuihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang 等NeurIPS 2025 · 被引用 50 次
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation ControlZekai Gu, Rui Yan, Jiahao Lu, Peng Li 等SIGGRAPH 2025 · 被引用 21 次
- BulletTime: Decoupled Control of Time and Camera Pose for Video GenerationYiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu 等CVPR 2026 · 被引用 15 次
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel 等ICLR 2026 · 被引用 13 次
它引用的顶会 Paper51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Zero-Shot Controllable Image-to-Video Animation via Motion DecompositionShoubin Yu, Jacob Zhiyuan Fang, Jian Zheng, Gunnar A. Sigurdsson 等ACM MM 2024 · 被引用 3 次
- IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot MannerYuyang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 等CVPR 2025
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang 等AAAI 2025 · 被引用 1 次
- TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsHaomiao Ni, Bernhard Egger, Suhas Lohit, Anoop Cherian 等CVPR 2024 · 被引用 7 次
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian 等SIGGRAPH 2024 · 被引用 66 次
