Anchor Frame Bridging for Coherent First-Last Frame Video Generation
Xuehan Hou, Meng Fan, Pengchong Qiao, Rat Cheng, Yian Zhao, Lei Zhu, Kaiwen Cheng, Chang Liu, Jie Chen
摘要
First-last frame video generation has recently gained significant attention. It enables coherent motion generation between specified first and last frames. However, this approach suffers from semantic degradation in intermediate frames, causing scene distortion and subject deformation that undermine temporal consistency. To address this issue, we introduce Anchor Frame Bridging (AFB), a novel plug-and-play method that explicitly bridges semantic continuity from boundary frames to intermediate frames, offering training-free adaptability and generalizability. By adaptively interpolating anchor frames at temporally critical locations exhibiting maximal semantic discontinuities, our approach effectively mitigates semantic drift in intermediate frames. Specifically, we propose an adaptive anchor frame selection module, which generates text-aligned candidate frames via frame order reversal and selects anchors based on semantic continuity. Subsequently, we develop anchor frame guided generation, which leverages the selected anchor frames to guide semantic propagation across intermediate frames, ensuring consistent boundary semantics and preserving temporal coherence throughout the video sequence. The final video is synthesized using the first frame, last frame, selected anchor frames, and the text prompt. The results demonstrate that our method significantly enhances the temporal consistency and overall quality of generated videos. Specifically, when applied to the Wan2.1-I2V model, it yields improvements of 16.58% in FVD and 10.21% in PSNR. The codes are provided in the supplementary material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 被引用 434 次
- ControlVideo: Training-free Controllable Text-to-video GenerationYabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang 等ICLR 2024 · 被引用 359 次
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 等ICLR 2023 · 被引用 313 次
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui 等ACM MM 2022 · 被引用 184 次
相关 Paper
- Anchoring and Rescaling Attention for Semantically Coherent InbetweeningTae Eun Choi, Sumin Shim, Junhyeok Kim, Seong Jae HwangCVPR 2026 · 被引用 2 次
- I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsXun Guo, Mingwu Zheng, Liang Hou, Yuan Gao 等SIGGRAPH 2024 · 被引用 26 次
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi 等KDD 2026
- Optical-Flow Guided Prompt Optimization for Coherent Video GenerationHyelin Nam, Jaemin Kim, Dohun Lee, Jong Chul YeCVPR 2025
- FrameBridge: Improving Image-to-Video Generation with Bridge ModelsYuji Wang, Zehua Chen, Xiaoyu Chen, Yixiang Wei 等ICML 2025
