FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing
Guangzhao Li, Yanming Yang, Chenxi Song, Xiaohong Liu, Chi Zhang
Abstract
Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm maps the video to a latent space before editing. However, the inversion process is not perfectly accurate, often compromising appearance fidelity and motion consistency.To address this, we introduce FlowDirector, a novel training-free and inversion-free video editing framework. Our framework models the editing process as a direct evolution in the data space. It guides the video to transition smoothly along its inherent spatio-temporal manifold using an ordinary differential equation (ODE), thereby avoiding the inaccurate inversion step. From this foundation, we introduce three flow correction strategies for appearance, motion, and stability: 1) Direction-aware flow correction amplifies components that oppose the source direction and removes irrelevant terms, breaking conservative streamlines and enabling stronger structural and textural changes. 2) Motion–appearance decoupling optimizes motion agreement as an energy term at each timestep, significantly improving consistency and motion transfer. 3) Differential averaging guidance strategy leverages differences among multiple candidate flows to approximate a low variance regime at low cost, suppressing artifacts and stabilizing the trajectory. Extensive experiments across various editing tasks and benchmarks demonstrate that FlowDirector achieves state-of-the-art performance in instruction following, temporal consistency, and background preservation, establishing an efficient new paradigm for coherent video editing without inversion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 223fe813-f7df-4e1b-8045-b68bae1dcd6aCited by top-tier papers5
- Lynx: Towards High-Fidelity Personalized Video GenerationShen Sang, Tiancheng Zhi, Tianpei Gu, Jing Liu et al.CVPR 2026 · 9 citations
- Free Lunch for Stabilizing Rectified Flow InversionChenru Wang, Beier Zhu, Chi ZhangICLR 2026 · 6 citations
- Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time TrainingRuicheng Zhang, Jun Zhou, Zunnan Xu, Zihao Liu et al.AAAI 2026 · 4 citations
- FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image EditingKaixiang Yang, Boyang Shen, Xin Li, Yuchen Dai et al.AAAI 2026 · 3 citations
- In-Context Generation with Regional Constraints for Instructional Video EditingZhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu et al.ICML 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu et al.CVPR 2025
- FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image EditingJeongsol Kim, Yeobin Hong, Jonghyun Park, Jong Chul YeICLR 2026 · 35 citations
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 22 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi et al.KDD 2026
