Lune

ACM MM2021顶会

Intrinsic Temporal Regularization for High-resolution Human Video Synthesis

Lingbo Yang, Zhanning Gao, Siwei Ma, Wen Gao

2021年份
1被引次数

摘要

Fashion video synthesis has attracted increasing attention due to its huge potential in immersive media, virtual reality and online retail applications, yet traditional 3D graphic pipelines often require extensive manual labor on data capture and model rigging. In this paper, we investigate an image-based approach to this problem that generates a fashion video clip from a still source image of the desired outfit, which is then rigged in a framewise fashion under the guidance of a driving video. A key challenge for this task lies in the modeling of feature transformation across source and driving frames, where fine-grained transform helps promote visual details at garment regions, but often at the expense of intensified temporal flickering. To resolve this dilemma, we propose a novel framework with 1) a multi-scale transform estimation and feature fusion module to preserve fine-grained garment details, and 2) an intrinsic regularization loss to enforce temporal consistency of learned transform between adjacent frames. Our solution is capable of generating 512512 fashion videos with rich garment details and smooth fabric movements beyond existing results. Extensive experiments over the FashionVideo benchmark dataset have demonstrated the superiority of the proposed framework over several competitive baselines.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖