High-Fidelity Virtual Try-On beyond Paired Data Scarcity via Diffusion-based Cycle-Consistent Learning
Jia Wu, Yijing Dai, Tingfeng Cao, Meiling Wu, Tao Luo, Jian Dong Zhang, Guangming Lu, Xiaoyi Zeng
Abstract
Diffusion-based virtual try-on methods rely on vast high-quality garment-person pairs, which are scarce in practice due to the high cost of data collection and preprocessing, limiting their performance in real-world scenarios.To overcome this bottleneck, we propose Cycle-Consistent Virtual Try-On (CCVTON), a diffusion-based approach that enables effective training using massive in-the-wild person images. Specifically, CCVTON introduces a Cycle-Consistent Learning (CCL) strategy that just employs a single unified generative model to disentangle a garment from a person image (try-off branch) and transfer it to the same individual (try-on branch), forming a reconstruction cycle. To this end, we first warm up a Unified Diffusion Transformer (UDiT) on open-source paired data to acquire basic try-on and try-off capabilities. When adapting UDiT to in-the-wild person images, we employ a Multi-Criteria Filtering Operation to select high-quality garments disentangled from person images by the pretrained UDiT. These filtered garments are not used as inputs for CCL, but serve as soft constraints for a perceptual regularization loss, preventing the try-off branch from collapsing to trivial copying. In addition, we propose a garment-aware mask generation with a two-stage refinement process to suppress garment leakage while maintaining person consistency.Extensive experiments show that CCVTON achieves state-of-the-art results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 568bf08f-1da8-4515-a993-692e57435e99Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 177 citations
- IMAGDressing-v1: Customizable Virtual DressingFei Shen, Xin Jiang, Xin He, Hu Ye et al.AAAI 2025 · 128 citations
Related papers
- Disentangled Cycle Consistency for Highly-Realistic Virtual Try-OnChongjian Ge, Yibing Song, Yuying Ge, Han Yang et al.CVPR 2021
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park et al.CVPR 2024
- Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationChenghu Du, Junyin Wang, Feng Yu, Shengwu XiongAAAI 2025 · 8 citations
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
- BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data TrainingXuanpu Zhang, Dan Song, Pengxin Zhan, Tianyu Chang et al.CVPR 2025
