Learning Fine-Grained Motion Embedding for Landscape Animation
Hongwei Xue, Bei Liu, Huan Yang, Jianlong Fu, Houqiang Li, Jiebo Luo
Abstract
In this paper we focus on landscape animation, which aims to generate time-lapse videos from a single landscape image. Motion is crucial for landscape animation as it determines how objects move in videos. Existing methods are able to generate appealing videos by learning motion from real time-lapse videos. However, current methods suffer from inaccurate motion generation, which leads to unrealistic video results. To tackle this problem, we propose a model named FGLA to generate high-quality and realistic videos by learning Fine-Grained motion embedding for Landscape Animation. Our model consists of two parts: (1) a motion encoder which embeds time-lapse motion in a fine-grained way. (2) a motion generator which generates realistic motion to animate input images. To train and evaluate on diverse time-lapse videos, we build the largest high-resolution Time-lapse video dataset with Diverse scenes, namely Time-lapse-D, which includes 16,874 video clips with over 10 million frames. Quantitative and qualitative experimental results demonstrate the superiority of our method. In particular, our method achieves relative improvements by 19% on LIPIS and 5.6% on FVD compared with state-of-the-art methods on our dataset. A user study carried out with 700 human subjects shows that our approach visually outperforms existing methods by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Improving Visual Quality of Image Synthesis by A Token-based Generator with TransformersYanhong Zeng, Huan Yang, Hongyang Chao, Jianbo Wang et al.NeurIPS 2021 · 31 citations
- TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input ViewsHsiang-Hui Hung, Huu-Phu Do, Yung-Hui Li, Ching-Chun HuangACM MM 2024 · 1 citation
- MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video GenerationLudan Ruan, Yiyang Ma, Huan Yang, Huiguo He et al.CVPR 2023
Builds on3
- Aesthetic-Aware Image Style TransferZhiyuan Hu, Jia Jia, Bei Liu, Yaohua Bu et al.ACM MM 2020 · 39 citations
- Time Flies: Animating a Still Image With Time-Lapse Video As ReferenceChia-Chi Cheng, Hung-Yu Chen, Wei-Chen ChiuCVPR 2020
- Learning Texture Transformer Network for Image Super-ResolutionFuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu et al.CVPR 2020
Related papers
- Aesthetics-Driven Virtual Time-Lapse Photography GenerationLihua Lu, Hui Wei, Xin Jin, Yihao Zhang et al.ACM MM 2023 · 1 citation
- Optimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImagesIn-Hwan Jin, Haesoo Choo, Seong-Hun Jeong, Park Heemoon et al.ICLR 2025
- Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single ImageLiao Shen, Xingyi Li, Huiqiang Sun, Juewen Peng et al.ACM MM 2023 · 15 citations
- Speech Driven Tongue AnimationSalvador Medina, Denis Tomè, Carsten Stoll, Mark Tiede et al.CVPR 2022 · 14 citations
- Animating General Image with Large Visual Motion ModelDengsheng Chen, Xiaoming Wei, Xiaolin WeiCVPR 2024
