UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition Framework
Rui Jiang, Chongmian Wang, Xinghe Fu, Yehao Lu, Teng Li, Xi Li
摘要
Video transitions are critical for ensuring temporal coherence in edited media, yet existing methods often rely on handcrafted effects or relative-scale trajectories that fail to capture the physical structure of real-world scenes. In this work, we introduce a scale-aware video transition framework that explicitly incorporates depth-aware 3D reasoning into a diffusion-based generation pipeline. Built upon a powerful I2V foundation, our method leverages single-image depth prediction to align camera motion with metric-scale geometry, enabling physically consistent transitions. To reduce reliance on precise camera inputs, we propose a bidirectional conditional control module and a progressive training strategy with conditional dropout, enhancing generalization to loosely specified or missing camera trajectories. Extensive experiments demonstrate that our approach achieves state-of-the-art performance, delivering realistic, geometrically coherent transitions across diverse scenes and applications with minimal input guidance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
- Preserve Your Own Correlation: A Noise Prior for Video Diffusion ModelsSongwei Ge, Seungjun Nah, Guilin Liu, Tyler Poon 等ICCV 2023 · 被引用 319 次
- SEINE: Short-to-Long Video Diffusion Model for Generative Transition and PredictionXinyuan Chen, Yaohui Wang, Lingjun Zhang, Shaobin Zhuang 等ICLR 2024 · 被引用 226 次
- XVFI: eXtreme Video Frame InterpolationHyeonjun Sim, Jihyong Oh, Munchurl KimICCV 2021 · 被引用 207 次
相关 Paper
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang 等NeurIPS 2025 · 被引用 18 次
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlTeng Li, Guangcong Zheng, Rui Jiang, Shuigen Zhan 等ICCV 2025 · 被引用 5 次
- Structure and Content-Guided Video Synthesis with Diffusion ModelsPatrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog 等ICCV 2023 · 被引用 733 次
- DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationYuejiang Dong, Wang Zhao, Jiale Xu, Ying Shan 等ICCV 2025
- I2V3D: Controllable Image-to-Video Generation with 3D GuidanceZhiyuan Zhang, Dongdong Chen, Jing LiaoICCV 2025 · 被引用 3 次
