Implicit Warping for Animation with Image Sets
Arun Mallya, Ting-Chun Wang, Ming-Yu Liu
Abstract
We present a new implicit warping framework for image animation using sets of source images through the transfer of the motion of a driving video. A single cross- modal attention layer is used to find correspondences between the source images and the driving image, choose the most appropriate features from different source images, and warp the selected features. This is in contrast to the existing methods that use explicit flow-based warping, which is designed for animation using a single source and does not extend well to multiple sources. The pick-and-choose capability of our framework helps it achieve state-of-the-art results on multiple datasets for image animation using both single and multiple source images. The project website is available at https://deepimagination.cc/implicit warping/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionJohanna Suvi Karras, Aleksander Holynski, Ting-Chun Wang, Ira Kemelmacher-ShlizermanICCV 2023 · 224 citations
- MagicAnimate: Temporally Consistent Human Image Animation using Diffusion ModelZhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan et al.CVPR 2024 · 106 citations
- SPACE: Speech-driven Portrait Animation with Controllable ExpressionSiddharth Gururani, Arun Mallya, Ting-Chun Wang, Rafael Valle et al.ICCV 2023 · 58 citations
- StyleAvatar: Real-time Photo-realistic Portrait Avatar from a Single VideoLizhen Wang, Xiaochen Zhao, Jingxiang Sun, Yuxiang Zhang et al.SIGGRAPH 2023 · 49 citations
- ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion ModelingQuanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu et al.NeurIPS 2024 · 21 citations
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
Related papers
- Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion TransferLi Yuze, Dong Gong, Xiao Cao, Junchao Yuan et al.CVPR 2026 · 3 citations
- Image Animation with Perturbed MasksYoav Shalev, Lior WolfCVPR 2022 · 6 citations
- Thin-Plate Spline Motion Model for Image AnimationJian Zhao, Hui ZhangCVPR 2022 · 196 citations
- MotionFlow: Attention-Driven Motion Transfer in Video Diffusion ModelsTuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar YanardagAAAI 2026
- JAFPro: Joint Appearance Fusion and Propagation for Human Video Motion Transfer from Multiple Reference ImagesXianggang Yu, Haolin Liu, Xiaoguang Han, Zhen Li et al.ACM MM 2020 · 1 citation
