Re-Attentional Controllable Video Diffusion Editing
Yuanzhi Wang, Yong Li, Mengyi Liu, Xiaoya Zhang, Xin Liu, Zhen Cui, Antoni B. Chan
摘要
Editing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-toimage diffusion models for text-guided video editing, resulting in remarkable video editing capabilities. However, they may still suffer from some limitations such as mislocated objects, incorrect number of objects. Therefore, the controllability of video editing remains a formidable challenge. In this paper, we aim to challenge the above limitations by proposing a Re-Attentional Controllable Video Diffusion Editing (ReAtCo) method. Specially, to align the spatial placement of the target objects with the edited text prompt in a trainingfree manner, we propose a Re-Attentional Diffusion (RAD) to refocus the cross-attention activation responses between the edited text prompt and the target video during the denoising stage, resulting in a spatially location-aligned and semantically high-fidelity manipulated video. In particular, to faithfully preserve the invariant region content with less border artifacts, we propose an Invariant Region-guided Joint Sampling (IRJS) strategy to mitigate the intrinsic sampling errors w.r.t the invariant regions at each denoising timestep and constrain the generated content to be harmonized with the invariant region content. Experimental results verify that ReAtCo consistently improves the controllability of video diffusion editing and achieves superior video editing performance. Codes are released at https://github.com/mdswyz/ReAtCo
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- FlowDirector: Training-Free Flow Steering for Precise Text-to-Video EditingGuangzhao Li, Yanming Yang, Chenxi Song, Xiaohong Liu 等CVPR 2026 · 被引用 27 次
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li 等CVPR 2026 · 被引用 5 次
- Value Diffusion Reinforcement LearningXiaoliang Hu, Fuyun Wang, Tong Zhang, Zhen CuiNeurIPS 2025 · 被引用 2 次
- Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionFuyun Wang, Tong Zhang, Yuanzhi Wang, Yide Qiu 等CVPR 2025
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationShuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu 等CVPR 2025
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- AVID: Any-Length Video Inpainting with Diffusion ModelZhixing Zhang, Bichen Wu, Xiaoyan Wang, Yaqiao Luo 等CVPR 2024 · 被引用 25 次
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 等ICLR 2024 · 被引用 175 次
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu 等CVPR 2025
- Video-P2P: Video Editing with Cross-Attention ControlShaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin 等CVPR 2024 · 被引用 99 次
- CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and CompatibilityBojia Zi, Shihao Zhao, Xianbiao Qi, Jianan Wang 等AAAI 2025 · 被引用 6 次
