GPD-VVTO: Preserving Garment Details in Video Virtual Try-On
Yuanbin Wang, Weilun Dai, Long Chan, Huanyu Zhou, Aixi Zhang, Si Liu
摘要
Video Virtual Try-On aims to transfer a garment onto a person in the video. Previous methods typically focus on image-based virtual try-on, but directly applying these methods to videos often leads to temporal discontinuity due to inconsistencies between frames. Limited attempts in video virtual try-on also suffer from unrealistic results and poor generalization ability. In light of previous research, we posit that the task of video virtual try-on can be decomposed into two key aspects: (1) single-frame results are realistic and natural, while retaining consistency with the garment; (2) the person's actions and the garment are coherent throughout the entire video. To address these two aspects, we propose a novel two-stage framework based on Latent Diffusion Model, namely Garment-Preserving Diffusion for Video Virtual Try-On (GPD-VVTO). In the first stage, the model is trained on single-frame data to improve the ability of generating high-quality try-on images. We integrate both low-level texture features and high-level semantic features of the garment into the denoising network to preserve garment details while ensuring a natural fit between the garment and the person. In the second stage, the model is trained on video data to enhance temporal consistency. We devise a novel Garment-aware Temporal Attention (GTA) module that incorporates garment features into temporal attention, enabling the model to maintain the fidelity to the garment during temporal modeling. Furthermore, we collect a video virtual try-on dataset containing high-resolution videos from diverse scenes, addressing the limited variety of current datasets in terms of video background and human actions. Extensive experiments demonstrate that our method outperforms existing state-of-the-art methods in both image-based and video-based virtual try-on tasks, indicating the effectiveness of our proposed framework.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- 3DV-TON: Textured 3D-Guided Consistent Video Try-on via Diffusion ModelsMin Wei, Chaohui Yu, Jingkai Zhou, Fan WangACM MM 2025 · 被引用 1 次
- Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single ImageJunkun Chen, Aayush Bansal, Minh Vo, Yu-Xiong WangNeurIPS 2025 · 被引用 1 次
- iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic GuidanceJun Zheng, Zhengze Xu, Mengting Chen, Chen Wenyin 等ICML 2026 · 被引用 1 次
- The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details InjectionQingdong He, Xueqin Chen, Yanjie Pan, Peng Tang 等CVPR 2026
相关 Paper
- Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose InteractionDong Li, Wenqi Zhong, Wei Yu, Yingwei Pan 等CVPR 2025
- MV-TON: Memory-based Video Virtual Try-on networkXiaojing Zhong, Zhonghua Wu, Taizhe Tan, Guosheng Lin 等ACM MM 2021 · 被引用 27 次
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang 等CVPR 2024 · 被引用 25 次
- SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion ModelsHung Nguyen, Quang Qui-Vinh Nguyen, Khoi Nguyen, Rang NguyenAAAI 2025 · 被引用 13 次
- FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-OnHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu 等ICCV 2019 · 被引用 130 次
