Anchor Frame Bridging for Coherent First-Last Frame Video Generation
Xuehan Hou, Meng Fan, Pengchong Qiao, Rat Cheng, Yian Zhao, Lei Zhu, Kaiwen Cheng, Chang Liu, Jie Chen
Abstract
First-last frame video generation has recently gained significant attention. It enables coherent motion generation between specified first and last frames. However, this approach suffers from semantic degradation in intermediate frames, causing scene distortion and subject deformation that undermine temporal consistency. To address this issue, we introduce Anchor Frame Bridging (AFB), a novel plug-and-play method that explicitly bridges semantic continuity from boundary frames to intermediate frames, offering training-free adaptability and generalizability. By adaptively interpolating anchor frames at temporally critical locations exhibiting maximal semantic discontinuities, our approach effectively mitigates semantic drift in intermediate frames. Specifically, we propose an adaptive anchor frame selection module, which generates text-aligned candidate frames via frame order reversal and selects anchors based on semantic continuity. Subsequently, we develop anchor frame guided generation, which leverages the selected anchor frames to guide semantic propagation across intermediate frames, ensuring consistent boundary semantics and preserving temporal coherence throughout the video sequence. The final video is synthesized using the first frame, last frame, selected anchor frames, and the text prompt. The results demonstrate that our method significantly enhances the temporal consistency and overall quality of generated videos. Specifically, when applied to the Wan2.1-I2V model, it yields improvements of 16.58% in FVD and 10.21% in PSNR. The codes are provided in the supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
- ControlVideo: Training-free Controllable Text-to-video GenerationYabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang et al.ICLR 2024 · 359 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- DiT: Self-supervised Pre-training for Document Image TransformerJunlong Li, Yiheng Xu, Tengchao Lv, Lei Cui et al.ACM MM 2022 · 184 citations
Related papers
- Anchoring and Rescaling Attention for Semantically Coherent InbetweeningTae Eun Choi, Sumin Shim, Junhyeok Kim, Seong Jae HwangCVPR 2026 · 2 citations
- I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsXun Guo, Mingwu Zheng, Liang Hou, Yuan Gao et al.SIGGRAPH 2024 · 26 citations
- VIVID: Backbone Training-Free Text-to-Image Video Editing via Variational Latent AnchorsZhangkai Wu, Xuhui Fan, Zhongyuan Xie, Kaize Shi et al.KDD 2026
- Optical-Flow Guided Prompt Optimization for Coherent Video GenerationHyelin Nam, Jaemin Kim, Dohun Lee, Jong Chul YeCVPR 2025
- FrameBridge: Improving Image-to-Video Generation with Bridge ModelsYuji Wang, Zehua Chen, Xiaoyu Chen, Yixiang Wei et al.ICML 2025
