Intrinsic Temporal Regularization for High-resolution Human Video Synthesis
Lingbo Yang, Zhanning Gao, Siwei Ma, Wen Gao
Abstract
Fashion video synthesis has attracted increasing attention due to its huge potential in immersive media, virtual reality and online retail applications, yet traditional 3D graphic pipelines often require extensive manual labor on data capture and model rigging. In this paper, we investigate an image-based approach to this problem that generates a fashion video clip from a still source image of the desired outfit, which is then rigged in a framewise fashion under the guidance of a driving video. A key challenge for this task lies in the modeling of feature transformation across source and driving frames, where fine-grained transform helps promote visual details at garment regions, but often at the expense of intensified temporal flickering. To resolve this dilemma, we propose a novel framework with 1) a multi-scale transform estimation and feature fusion module to preserve fine-grained garment details, and 2) an intrinsic regularization loss to enforce temporal consistency of learned transform between adjacent frames. Our solution is capable of generating 512512 fashion videos with rich garment details and smooth fabric movements beyond existing results. Extensive experiments over the FashionVideo benchmark dataset have demonstrated the superiority of the proposed framework over several competitive baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7443de57-4248-409e-ba90-c84025ef99e7Builds on9
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View SynthesisWen Liu, Zhixin Piao, Jie Min, Wenhan Luo et al.ICCV 2019 · 285 citations
- Blind Video Temporal Consistency via Deep Video PriorChenyang Lei, Yazhou Xing, Qifeng ChenNeurIPS 2020 · 134 citations
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu et al.ICCV 2019 · 130 citations
Related papers
- DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionJohanna Suvi Karras, Aleksander Holynski, Ting-Chun Wang, Ira Kemelmacher-ShlizermanICCV 2023 · 224 citations
- Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet SupervisionHyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul JooCVPR 2026 · 1 citation
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou et al.ACM MM 2024 · 4 citations
- ZFlow: Gated Appearance Flow-based Virtual Try-on with 3D PriorsAyush Chopra, Rishabh Jain, Mayur Hemani, Balaji KrishnamurthyICCV 2021 · 73 citations
- Robust-MVTON: Learning Cross-Pose Feature Alignment and Fusion for Robust Multi-View Virtual Try-OnNannan Zhang, Yijiang Li, Dong Du, Zheng Chong et al.CVPR 2025
