MotiF: Making Text Count in Image Animation with Motion Focal Loss
Shijie Wang, Samaneh Azadi, Rohit Girdhar, Saketh Rambhatla, Chen Sun, Xi Yin
摘要
Baseline (a) (b) (c) ⊕ L2 loss Condition leakage "Common vole drags sunflower seeds in the hole." Re-weighting A bear jumping high on a meadow. A bear running to the left on a meadow. Figure 1. Motivation and results of MotiF. (a) Example video frames and the corresponding motion heatmaps calculated from optical flow. In this example, 97% of the pixels are static while only 3% has meaningful motion. (b) In standard TI2V training pipeline, the model may learn to over-rely on the conditional image to optimize the L2 loss. This issue has been identified in [53] and termed as conditional image leakage. We propose MotiF to guide the model's learning to focus on regions with more motion via motion heatmap re-weighting. (c) Qualitative results comparing MotiF to the baseline on examples from our proposed TI2V-Bench evaluation set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ReSim: Reliable World Simulation for Autonomous DrivingJiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 等NeurIPS 2025 · 被引用 53 次
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang 等NeurIPS 2025 · 被引用 18 次
- WorldReel: 4D Video Generation with Consistent Geometry and Motion ModelingShaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra 等CVPR 2026 · 被引用 3 次
- VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video ModelsHila Chefer, Uriel Singer, Amit Zohar, Yuval Kirstain 等ICML 2025
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- MoCha-Stereo: Motif Channel Attention Network for Stereo MatchingZiyang Chen, Wei Long, He Yao, Yongjun Zhang 等CVPR 2024
- Motion Attribution for Video GenerationXindi Wu, Despoina Paschalidou, Jun Gao, Antonio Torralba 等ICML 2026
- Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion ModelMin Zhao, Hongzhou Zhu, Chendong Xiang, Kaiwen Zheng 等NeurIPS 2024 · 被引用 33 次
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel 等ICLR 2026 · 被引用 13 次
- Streaming Video ModelYucheng Zhao, Chong Luo, Chuanxin Tang, Dongdong Chen 等CVPR 2023
