VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
Xiangpeng Yang, Linchao Zhu, Hehe Fan, Yi Yang
摘要
Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a formidable challenge. The major difficulties in multi-grained editing include semantic misalignment of text-to-region control and feature coupling within the diffusion model. To address these difficulties, we present VideoGrain, a zeroshot approach that modulates space-time (cross-and self-) attention mechanisms to achieve fine-grained control over video content. We enhance text-to-region control by amplifying each local prompt's attention to its corresponding spatialdisentangled region while minimizing interactions with irrelevant areas in crossattention. Additionally, we improve feature separation by increasing intra-region awareness and reducing inter-region interference in self-attention. Extensive experiments demonstrate our method achieves state-of-the-art performance in realworld scenarios. Our code, data, and demos are available on the project page. While existing methods employ various visual consistency techniques, such as optical flow (Cong et al., 2023;Yang et al., 2023), control signals (Zhang et al., 2023b), or feature correspondence (Geyer et al., 2023). These methods remain instance-agnostic, often mixing features of different instances during editing (see Fig. 2 right). Ground-A-Video (Jeong & Ye, 2023), which inherits text-to-bounding box generation priors (Li et al., 2023), should be instance-level editing but still suffer from artifacts. Similarly, recent T2V-based methods like DMT (Yatim et al., 2024) and Pika (pik), although equipped with video generation priors, struggle with multi-grained edits. We find that the core issue is that diffusion models tend to treat different instances as the same class segments, leading to strong feature coupling across instances, as illustrated in Figure 3.
To address this problem, our primary insight is to 1) enable text-to-region control and 2) keep feature separation between regions. In the typical diffusion models, the cross-attention layer serves as a key component to update textual features control over each spatial region, while the self-attention layer generates globally coherent structures by connecting each frame token across time. Therefore, we propose Spatial-Temporal Layout-Guided Attention (ST-Layout Attn), which modulates both spacetime cross-and self-attention in a unified manner to achieve the above goals.
In the cross-attention layer, the uniform application of global text prompts across all frame tokens leads to severe semantic misalignment, which reduces the precision of multi-grained text-to-region control. To address this, we modulate cross-attention to amplify each local prompt's focus on its corresponding spatial-disentangled region while suppressing attention to irrelevant areas. In the self-attention layer, pixels from one region may attend to outside or similar regions within the same class, leading to feature coupling and texture mixing, which is an inherent limitation of diffusion models that complicates multi-grained video editing. To mitigate this, we modulate self-attention to enhance feature separation by increasing intra-region focus and reducing inter-region interactions, ensuring each query attends only to its target region.
Our key contributions can be summarized as follows:
• To the best of our knowledge, this is the first attempt at multi-grained video editing. Our method enables both class-level, instance-level and part-level editing.
• We propose a novel framework, dubbed VideoGrain, which modulates spatial-temporal cross-and self-attention for text-to-region control and feature separation between regions.
• Without tuning any parameters, we achieve state-of-the-art results on existing benchmarks and real-world videos both qualitatively and quantitatively.
2 RELATED WORK
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic DatasetQingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu 等CVPR 2026 · 被引用 79 次
- EffiVMT: Video Motion Transfer via Efficient Spatial-Temporal Decoupled FinetuningYue Ma, Yulong Liu, Qiyuan Zhu, Xiangpeng Yang 等ICLR 2026 · 被引用 70 次
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 被引用 7 次
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma 等CVPR 2026 · 被引用 6 次
- EEdit ⚡: Rethinking the Spatial and Temporal Redundancy for Efficient Image EditingZexuan Yan, Yue Ma, Chang Zou, Wenteng Chen 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper38
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao 等NeurIPS 2024 · 被引用 76 次
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 等ICLR 2024 · 被引用 175 次
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu 等CVPR 2025
- FateZero: Fusing Attentions for Zero-shot Text-based Video EditingChenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei 等ICCV 2023 · 被引用 510 次
- VidToMe: Video Token Merging for Zero-Shot Video EditingXirui Li, Chao Ma, Xiaokang Yang, Ming-Hsuan YangCVPR 2024
