Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image
Junkun Chen, Aayush Bansal, Minh Vo, Yu-Xiong Wang
Abstract
We introduce the Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and lengthy video data, while providing the flexibility to generate videos of arbitrary length. The key challenges of this task are twofold: ensuring local smoothness between adjacent segments and maintaining global temporal consistency across different segments. To address these challenges, we propose our VFR framework, which ensures smoothness through a prefix video condition and enforces consistency with the anchor video-a 360 • video that comprehensively captures the human's wholebody appearance. Our VFR generates minute-scale virtual try-on videos with both local smoothness and global temporal consistency under various motions, making it a pioneering work in long virtual try-on video generation. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 722984f1-1c4e-4797-bfb9-29d4385ed4b9Builds on48
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
Related papers
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu et al.ICCV 2019 · 130 citations
- Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in VideosZhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing et al.ACM MM 2024 · 14 citations
- ClothFormer: Taming Video Virtual Try-on in All ModuleJianbin Jiang, Tan Wang, He Yan, Junhui LiuCVPR 2022 · 29 citations
- SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion ModelsHung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, Rang NguyenAAAI 2025 · 13 citations
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou et al.ACM MM 2024 · 4 citations
