Are Image-to-Video Models Good Zero-Shot Image Editors?
Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi Yang
Abstract
Large-scale video diffusion models exhibit strong world-simulation and temporal reasoning capabilities, yet their potential as zero-shot image editors remains underexplored. We present IF-Edit (Image Edit by Generating Frames), a tuning-free framework that repurposes pre-trained image-to-video diffusion models for instruction-driven image editing. IF-Edit addresses three core obstacles—prompt misalignment, redundant temporal latents, and blurry late-stage frames—via: (1) a Chain-of-Thought Prompt Enhancement module that reformulates static editing instructions into temporally grounded reasoning prompts; (2) a Temporal Latent Dropout strategy that compresses frame latents after the expert-switch point, accelerating denoising while preserving global semantics and temporal coherence; and (3) a Self-Consistent Post-Refinement step that refines the sharpest late-stage frame through a brief still-video trajectory, leveraging the video prior for sharper and more faithful results. Extensive experiments across four public benchmarks—covering non-rigid deformations, physical and temporal reasoning, and general instruction editing—show that IF-Edit achieves strong performance on non-rigid and reasoning-centric tasks while remaining competitive on general-purpose edits. Our study offers a systematic view of video diffusion models as image editors, revealing their unique strengths, limitations, and a simple recipe for unified video–image generative reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae7d1b31-448a-4c16-b7c6-c0f84fd2b667Cited by top-tier papers3
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance GenerationRuihang Xu, Dewei Zhou, Fan Ma, Yi YangICLR 2026 · 19 citations
- CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image GenerationChengzhuo Tong, Chang Mingkun, Shenglong Zhang, Yuran Wang et al.ICML 2026 · 7 citations
- Progressive Photorealistic SimplificationAdi Rosenthal, Dana Berman, Yedid Hoshen, Ariel ShamirSIGGRAPH 2026
Builds on61
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma et al.CVPR 2026 · 6 citations
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao et al.NeurIPS 2024 · 76 citations
- Instruction-Based Image Editing with Planning, Reasoning, and GenerationLiya Ji, Chenyang Qi, Qifeng ChenICCV 2025 · 3 citations
- ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World SimulationJay Zhangjie Wu, Xuanchi Ren, Tianchang Shen, Tianshi Cao et al.ICLR 2026 · 17 citations
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi et al.KDD 2026
