All Parts Matter: A Unified Mask-Free Virtual Try-On Framework
Chenghu Du, Shengwu Xiong, Yi Rong
Abstract
Current virtual try-on methods primarily enhance performance through network optimization, employing strategies such as coarse-to-fine structures and ReferenceNet to inject clothing information. However, these methods are often constrained by the limited quantity and diversity of training samples, which ultimately restricts their potential for further improvement. To address this challenge, we propose a unified, mask-free virtual try-on framework. Our approach leverages the inherent strengths of latent diffusion models to enhance the capability of each pipeline component to accurately model the target distribution, thereby achieving superior performance. Specifically, we introduce a text-driven pseudo-input preparation strategy that significantly increases the diversity of clothing regions within the generated person pseudo-samples. This encourages the generator to focus on variations in these areas and improves the model's generalization capability. Within the generator itself, we develop a gated manipulation mechanism to prevent weight forgetting and reduce training costs. Furthermore, we incorporate a texture-aware injection module to explicitly integrate human-perceptible clothing texture information into the generation process. During inference, we propose a refining conditional inference strategy that mitigates the randomness introduced by Gaussian noise. This effectively preserves identity information and fine clothing details in the final results. Extensive experiments demonstrate that our method outperforms existing state-of-the-art virtual try-on approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21d34522-67a0-43bb-a0ea-049a7fcc2a7bBuilds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
- ClothFlow: A Flow-Based Model for Clothed Person GenerationXintong Han, Weilin Huang, Xiaojun Hu, Matthew R. ScottICCV 2019 · 297 citations
Related papers
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 177 citations
- Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label GenerationChenghu Du, Junyin Wang, Feng Yu, Shengwu XiongAAAI 2025 · 8 citations
- PG-VTON: Single-Pass Training-Free Virtual Try-On via Patch-Guided Reference AlignmentGuohao Zhao, Yuxin PengCVPR 2026
- BooW-VTON: Boosting In-the-Wild Virtual Try-On via Mask-Free Pseudo Data TrainingXuanpu Zhang, Dan Song, Pengxin Zhan, Tianyu Chang et al.CVPR 2025
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
