Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
Guy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin, Yaniv Taigman, Yossi Adi, Sagie Benaim, Adam Polyak
摘要
Abstract We consider the task of Image-to-Video (I2V) generation, which involves transforming static images into realistic video sequences based on a textual description. While recent advancements produce photorealistic outputs, they frequently struggle to create videos with accurate and consistent object motion, especially in multi-object scenarios. To address these limitations, we propose a two-stage compo-sitional framework that decomposes I2V generation into: (i) An explicit intermediate representation generation stage, followed by (ii) A video generation stage that is conditioned on this representation. Our key innovation is the introduction of a mask-based motion trajectory as an intermediate representation, that captures both semantic object information and motion, enabling an expressive but compact representation of motion and semantics. To incorporate the learned representation in the second stage, we utilize object-level attention objectives. Specifically, we consider a spatial, per-object, masked-cross attention objective, integrating object-specific prompts into corresponding latent space regions and a masked spatio-temporal self-attention objective, ensuring frame-to-frame consistency for each object. We evaluate our method on challenging benchmarks with multi-object and high-motion scenarios and empirically demonstrate that the proposed method achieves state-of-the-art results in temporal coherence, motion realism, and text-prompt faithfulness. Additionally, we introduce SA-V-128, a new challenging benchmark for singleobject and multi-object I2V generation, and demonstrate our method's superiority on this benchmark. Project page is available at https://guyyariv.github.io/TTM/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video GenerationJiehui Huang, Yuechen Zhang, Xu He, Yuan Gao 等CVPR 2026 · 被引用 12 次
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang 等ICCV 2025 · 被引用 10 次
- MAD: Motion Appearance Decoupling for efficient Driving World ModelsAhmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord 等CVPR 2026 · 被引用 7 次
- O-DisCo-Edit: Object Distortion Control for Unified Realistic Video EditingYuqing Chen, Junjie Wang, Lin Liu, Ruihang Chu 等AAAI 2026 · 被引用 6 次
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 被引用 6 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion ModelingXiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian 等SIGGRAPH 2024 · 被引用 66 次
- Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion TransferLi Yuze, Dong Gong, Xiao Cao, Junchao Yuan 等CVPR 2026 · 被引用 3 次
- MOSO: Decomposing MOtion, Scene and Object for Video PredictionMingzhen Sun, Weining Wang, Xinxin Zhu, Jing LiuCVPR 2023
- Comp-Attn: Present-and-Align Attention for Compositional Video GenerationHongyu Zhang, Yufan Deng, Shenghai Yuan, Xuehan Hou 等ICML 2026
- TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video GenerationXingrui Wang, Xin Li, Yaosi Hu, Hanxin Zhu 等AAAI 2025 · 被引用 3 次
