Arbitrary Generative Video Interpolation
Guozhen Zhang, Haiguang Wang, Chunyu Wang, Yuan Zhou, Qinglin Lu, Limin Wang
Abstract
Generative Video Frame Interpolation (VFI), which synthesizes intermediate frames from a given pair of start and end frames, plays a pivotal role in video creation. However, existing generative VFI methods are constrained to producing a fixed number of intermediate frames, which significantly limits the flexibility in adjusting the frame rate or duration of videos during the creation process. In this work, we present ArbInterp, a novel generative VFI framework that enables efficient interpolation at any timestamp and of any length. Specifically, to support interpolation at any timestamp, we propose the Timestamp-aware Rotary Position Embedding (TaRoPE), which modulates positions in temporal RoPE to align generated frames with target normalized timestamps. This design enables fine-grained control over frame timestamps, addressing the inflexibility of fixed-position paradigms in prior work. For any-length interpolation, we decompose long-sequence generation into segment-wise frame synthesis. We further design a novel appearance-motion decoupled conditioning strategy: it leverages prior segment endpoints to enforce appearance consistency and temporal semantics to maintain motion coherence, ensuring seamless spatiotemporal transitions across segments. Experimentally, we build comprehensive benchmarks for multi-scale frame interpolation (2× to 32×) to assess generalizability across arbitrary interpolation factors. Results show that ArbInterp outperforms prior methods across all scenarios with higher fidelity and more seamless spatiotemporal continuity. Video demos are provided on the website: https://mcg-nju.github.io/ArbInterp-Web.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and EditingDongxu Li, Junnan Li, Steven C. H. HoiNeurIPS 2023 · 587 citations
Related papers
- VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information AssumptionTianxiong Zhong, Xingye Tian, Boyuan Jiang, Xuebo Wang et al.NeurIPS 2025 · 4 citations
- Anchoring and Rescaling Attention for Semantically Coherent InbetweeningTae Eun Choi, Sumin Shim, Junhyeok Kim, Seong Jae HwangCVPR 2026 · 2 citations
- SpeedVFI: One-step Diffusion for Efficient Video Frame InterpolationGanggui Ding, Xiaogang Xu, Hao Chen, Chunhua ShenICML 2026
- ReRoPE: Repurposing RoPE for Relative Camera ControlChunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou et al.SIGGRAPH 2026 · 2 citations
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language ModelsHaoran Li, Yingjie Qin, Baoyuan Ou, Lai Xu et al.NeurIPS 2025 · 4 citations
