AlignVid: Taming Visual Dominance via Training-Free Attention Modulation in Text-guided Image-to-Video Generation
Yexin Liu, Wenjie Shu, Zile Huang, Haoze Zheng, Yueze Wang, Manyuan Zhang, Jinjing Zhu, Ser-Nam Lim, Harry Yang
摘要
Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial changes to a reference image (e.g., object addition, removal, or modification). Empirically, our analysis reveals that this stems from visual dominance, where the reference image causes severe attention dispersion, inhibiting the model's ability to incorporate new semantic information. To address this, we propose AlignVid, a training-free intervention that recalibrates the model's internal attention distribution. Drawing on an energy-based perspective of attention, AlignVid employs Attention Scaling Modulation (ASM) to reduce attention entropy and concentrate focus on semantic tokens, alongside Guidance Scheduling (GS) to maintain generation stability. To rigorously assess this capability, we present OmitI2V, a comprehensive benchmark for evaluating prompt adherence across object modification, addition, and deletion. Extensive experiments demonstrate that Align-Vid effectively enhances semantic fidelity with negligible computational overhead. Code and the OmitI2V benchmark are available at https: //github.com/LAW1223/AlignVid.
Image-path: OmitI2V/deletion/vanish/nature/2.jpg Prompt: The lush green mountain gradually erodes and disappears, leaving behind rolling sand dunes and a barren desert landscape.
Expected change: The green vegetation and rocky outcrops of the mountain fade away until only dunes of sand remain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and EditingMingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan 等ICCV 2023 · 被引用 770 次
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen 等NeurIPS 2023 · 被引用 579 次
- FateZero: Fusing Attentions for Zero-shot Text-based Video EditingChenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei 等ICCV 2023 · 被引用 510 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
相关 Paper
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass GuidanceJune Suk Choi, Kyungmin Lee, Sihyun Yu, Yisol Choi 等CVPR 2026 · 被引用 4 次
- Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual GenerationSubin Kim, Sangwoo Mo, Mamshad Nayeem Rizve, Yiran Xu 等CVPR 2026 · 被引用 2 次
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video GenerationMingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing 等CVPR 2026 · 被引用 8 次
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi 等KDD 2026
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min 等ACM MM 2025 · 被引用 1 次
