VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
Xiangpeng Yang, Linchao Zhu, Hehe Fan, Yi Yang
Abstract
Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a formidable challenge. The major difficulties in multi-grained editing include semantic misalignment of text-to-region control and feature coupling within the diffusion model. To address these difficulties, we present VideoGrain, a zeroshot approach that modulates space-time (cross-and self-) attention mechanisms to achieve fine-grained control over video content. We enhance text-to-region control by amplifying each local prompt's attention to its corresponding spatialdisentangled region while minimizing interactions with irrelevant areas in crossattention. Additionally, we improve feature separation by increasing intra-region awareness and reducing inter-region interference in self-attention. Extensive experiments demonstrate our method achieves state-of-the-art performance in realworld scenarios. Our code, data, and demos are available on the project page. While existing methods employ various visual consistency techniques, such as optical flow (Cong et al., 2023;Yang et al., 2023), control signals (Zhang et al., 2023b), or feature correspondence (Geyer et al., 2023). These methods remain instance-agnostic, often mixing features of different instances during editing (see Fig. 2 right). Ground-A-Video (Jeong & Ye, 2023), which inherits text-to-bounding box generation priors (Li et al., 2023), should be instance-level editing but still suffer from artifacts. Similarly, recent T2V-based methods like DMT (Yatim et al., 2024) and Pika (pik), although equipped with video generation priors, struggle with multi-grained edits. We find that the core issue is that diffusion models tend to treat different instances as the same class segments, leading to strong feature coupling across instances, as illustrated in Figure 3.
To address this problem, our primary insight is to 1) enable text-to-region control and 2) keep feature separation between regions. In the typical diffusion models, the cross-attention layer serves as a key component to update textual features control over each spatial region, while the self-attention layer generates globally coherent structures by connecting each frame token across time. Therefore, we propose Spatial-Temporal Layout-Guided Attention (ST-Layout Attn), which modulates both spacetime cross-and self-attention in a unified manner to achieve the above goals.
In the cross-attention layer, the uniform application of global text prompts across all frame tokens leads to severe semantic misalignment, which reduces the precision of multi-grained text-to-region control. To address this, we modulate cross-attention to amplify each local prompt's focus on its corresponding spatial-disentangled region while suppressing attention to irrelevant areas. In the self-attention layer, pixels from one region may attend to outside or similar regions within the same class, leading to feature coupling and texture mixing, which is an inherent limitation of diffusion models that complicates multi-grained video editing. To mitigate this, we modulate self-attention to enhance feature separation by increasing intra-region focus and reducing inter-region interactions, ensuring each query attends only to its target region.
Our key contributions can be summarized as follows:
• To the best of our knowledge, this is the first attempt at multi-grained video editing. Our method enables both class-level, instance-level and part-level editing.
• We propose a novel framework, dubbed VideoGrain, which modulates spatial-temporal cross-and self-attention for text-to-region control and feature separation between regions.
• Without tuning any parameters, we achieve state-of-the-art results on existing benchmarks and real-world videos both qualitatively and quantitatively.
2 RELATED WORK
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f8162fa-6c7a-4608-b242-bffaf4db5e50Cited by top-tier papers21
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic DatasetQingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu et al.CVPR 2026 · 79 citations
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang et al.ICLR 2026 · 70 citations
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 7 citations
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma et al.CVPR 2026 · 6 citations
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen et al.ICCV 2025 · 5 citations
Builds on38
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao et al.NeurIPS 2024 · 76 citations
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen et al.ICLR 2024 · 175 citations
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu et al.CVPR 2025
- FateZero: Fusing Attentions for Zero-shot Text-based Video EditingChenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei et al.ICCV 2023 · 510 citations
- VidToMe: Video Token Merging for Zero-Shot Video EditingXirui Li, Chao Ma, Xiaokang Yang, Ming-Hsuan YangCVPR 2024
