SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
Hung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, Rang Nguyen
Abstract
Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. While significant advances have been made in image-based virtual try-ons, extending these successes to video often results in frame-to-frame inconsistencies. Some approaches have attempted to address this by increasing the overlap of frames across multiple video chunks, but this comes at a steep computational cost due to the repeated processing of the same frames, especially for long video sequence. To address these challenges, we reconceptualize video virtual try-on as a conditional video inpainting task, with garments serving as input conditions. Specifically, our approach enhances image diffusion models by incorporating temporal attention layers to improve temporal coherence. To reduce computational overhead, we introduce ShiftCaching, a novel technique that maintains temporal consistency while minimizing redundant computations. Furthermore, we introduce the TikTokDress dataset, a new video tryon dataset featuring more complex backgrounds, challenging movements, and higher resolution compared to existing public datasets. Extensive experiments show that our approach outperforms current baselines, particularly in terms of video consistency and inference speed. Code and dataset will be made available upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet SupervisionHyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul JooCVPR 2026 · 1 citation
- The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details InjectionQingdong He, Xueqin Chen, Yanjie Pan, Peng Tang et al.CVPR 2026
- MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired DataHunor Laczkó, Libang Jia, Loc-Phat Truong, Diego Hernández et al.CVPR 2026
Builds on21
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
- Towards Multi-Pose Guided Virtual Try-On NetworkHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang et al.ICCV 2019 · 226 citations
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu et al.ICCV 2019 · 130 citations
Related papers
- GPD-VVTO: Preserving Garment Details in Video Virtual Try-OnYuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou et al.ACM MM 2024 · 4 citations
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion ModelsMin Wei, Chaohui Yu, Jingkai Zhou, Fan WangACM MM 2025 · 1 citation
- Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in VideosZhengze Xu, Mengting Chen, Zhao Wang, Linyu Xing et al.ACM MM 2024 · 14 citations
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionDong Li, Wenqi Zhong, Wei Yu, Yingwei Pan et al.CVPR 2025
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
