Motion-Residual Conflict-Aware Time Reversal for Generative Inbetweening
Zhenbang zhang, Zihui Cui, Haythem El-Messiry, Renmin Han, Zhiqiang Xu
Abstract
Image-to-video (I2V) diffusion models have recently made generative inbetweening a practical reality by synthesizing semantically plausible intermediate frames between two keyframes. Among them, inference-time sampling schemes that re-use large pre-trained I2V backbones without any additional training are especially attractive. Yet current methods frequently exhibit temporal inconsistency and artifacts such as ghosting or reverse motion. A key reason is that the two trajectories are driven by distinct motion priors, each inherited from its own conditioning frame, and are simply stitched together without explicitly reconciling these priors. We introduce Motion-Residual Conflict-Aware Time Reversal (MR-CATR), an inference-time sampling framework that aligns conflicting motion priors instead of discarding one of them or collapsing to a single start-conditioned prior. MR-CATR first derives a motion-residual-based direction from the forward path, combined with an end-conditioned residual to form a consensus motion axis. This design suppresses bidirectional motion conflicts while still allowing end-frame information to refine the trajectory and enforce endpoint consistency. MR-CATR can be seamlessly integrated into existing time-reversal samplers without changing model parameters. Experiments on generative inbetweening benchmarks show that our method produces videos with smoother motion, fewer artifacts, and consistently better quantitative scores and user preferences than prior strategies. Code is available at https://github.com/ zhangzhenbang2021/MR-CATR.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8295ba9f-666f-4091-9c3b-73d3607e08c4Builds on15
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu et al.CVPR 2022 · 128 citations
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian et al.SIGGRAPH 2024 · 66 citations
- VFIMamba: Video Frame Interpolation with State Space ModelsGuozhen Zhang, Chunxu Liu, Yutao Cui, Xiaotong Zhao et al.NeurIPS 2024 · 48 citations
- Generalizable Implicit Motion Modeling for Video Frame InterpolationZujin Guo, Wei Li, Chen Change LoyNeurIPS 2024 · 24 citations
Related papers
- Motion Prior Distillation in Time Reversal Sampling for Generative InbetweeningWooseok Jeon, Seunghyun Shin, Dongmin Shin, Hae-Gon JeonICLR 2026 · 5 citations
- Generative Inbetweening: Adapting Image-to-Video Models for Keyframe InterpolationXiaojuan Wang, Boyang Zhou, Brian Curless, Ira Kemelmacher-Shlizerman et al.ICLR 2025
- Trajectory-Stabilized Inference for Diffusion-Based Video InpaintingZhanhe Zhang, Jiahua Li, Xu Yang, Kun Wei et al.ICML 2026
- MotionRAG: Motion Retrieval-Augmented Image-to-Video GenerationChenhui Zhu, Yilu Wu, Shuai Wang, Gangshan Wu et al.NeurIPS 2025 · 8 citations
- Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You ThinkJie Tian, Xiaoye Qu, Zhenyi Lu, Wei Wei et al.CVPR 2025
