ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment
Yiyang Chen, Xuanhua He, Xiujun Ma, Jack Ma
Abstract
Training-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal consistency. Existing methods, often designed for U-Net architectures, suffer from two primary limitations: inaccurate inversion due to first-order solvers, and contextual conflicts caused by crude "hard" feature replacement. These issues are more challenging in Diffusion Transformers (DiTs), where the unsuitability of prior layer-selection heuristics makes effective guidance challenging. To address these limitations, we introduce ContextFlow, a novel training-free framework for DiT-based video object editing. In detail, we first employ a high-order Rectified Flow solver to establish a robust editing foundation. The core of our framework is Adaptive Context Enrichment (for specifying what to edit), a mechanism that addresses contextual conflicts. Instead of replacing features, it enriches the self-attention context by concatenating Key-Value pairs from parallel reconstruction and editing paths, empowering the model to dynamically fuse information. Additionally, to determine where to apply this enrichment (for specifying where to edit), we propose a systematic, data-driven analysis to identify task-specific vital layers. Based on a novel Guidance Responsiveness Metric, our method pinpoints the most influential DiT blocks for different tasks (e.g., insertion, swapping), enabling targeted and highly effective guidance. Extensive experiments show that ContextFlow significantly outperforms existing training-free methods and even surpasses several state-of-the-art training-based approaches, delivering temporally coherent, high-fidelity results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f7c4548-b822-4eef-8207-431db4f634bcCited by top-tier papers10
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang et al.ICLR 2026 · 70 citations
- FastVMT: Eliminating Redundancy in Video Motion TransferYue Ma, Zhikai Wang, Tianhao Ren, Mingzhe Zheng et al.ICLR 2026 · 32 citations
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region ControlZeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang et al.ICLR 2026 · 29 citations
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang et al.AAAI 2026 · 16 citations
- Group Editing: Edit Multiple Images in One GoYue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang et al.CVPR 2026 · 15 citations
Builds on60
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- FateZero: Fusing Attentions for Zero-shot Text-based Video EditingChenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei et al.ICCV 2023 · 510 citations
Related papers
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin et al.ICCV 2025 · 2 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
- Object-WIPER: Training-Free Object and Associated Effect Removal in VideosSaksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep KulkarniCVPR 2026 · 5 citations
- T-Edit: Triple-Branch Diffusion Anchoring for Consistent EditingLinsong Shan, Laurence Yang, Zecan Yang, Shijie Lian et al.ICML 2026
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi et al.KDD 2026
