MV-VTON: Multi-View Virtual Try-On with Diffusion Models
Haoyu Wang, Zhilu Zhang, Donglin Di, Shiliang Zhang, Wangmeng Zuo
Abstract
The goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the clothing and person are significantly inconsistent, particularly when the person's view is non-frontal, the results are unsatisfactory. To address this challenge, we introduce Multi-View Virtual Try-ON (MV-VTON), which aims to reconstruct the dressing results from multiple views using the given clothes. Given that single-view clothes provide insufficient information for MV-VTON, we instead employ two images, i.e., the frontal and back views of the clothing, to encompass the complete view as much as possible. Moreover, we adopt diffusion models that have demonstrated superior abilities to perform our MV-VTON. In particular, we propose a view-adaptive selection method where hard-selection and soft-selection are applied to the global and local clothing feature extraction, respectively. This ensures that the clothing features are roughly fit to the person's view. Subsequently, we suggest joint attention blocks to align and fuse clothing features with person features. Additionally, we collect a MV-VTON dataset MVG, in which each person has multiple photos with diverse views and poses. Experiments show that the proposed method not only achieves state-of-the-art results on MV-VTON task using our MVG dataset, but also has superiority on frontal-view virtual try-on task using VITON-HD and DressCode datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dfe734e-1f1c-4dbf-82d6-cbda64666fe3Cited by top-tier papers12
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang et al.CVPR 2026 · 59 citations
- MoSA: Motion-Coherent Human Video Generation via Structure-Appearance DecouplingHaoyu Wang, Hao Tang, Donglin Di, Zhilu Zhang et al.ICLR 2026 · 4 citations
- FashionComposer: Compositional Fashion Image GenerationSihui Ji, Yiyang Wang, Xi Chen, Xiaogang Xu et al.SIGGRAPH 2025 · 3 citations
- E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-ThoughtMeiqi Sun, Ming-Yu Liu, Junxiong ZhuCVPR 2026 · 2 citations
- UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic AlignmentWei Zhang, Yeying Jin, Xin Li, Yan Zhang et al.AAAI 2026 · 1 citation
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationYuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai et al.ICCV 2023 · 469 citations
Related papers
- Stable VITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-OnJeongho Kim, Gyojung Gu, Minho Park, Sunghyun Park et al.CVPR 2024
- MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-OnXiaoyu Han, Chenyang Wang, Jing Wang, Shunyuan Zheng et al.CVPR 2026
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
- MV-TON: Memory-based Video Virtual Try-on networkXiaojing Zhong, Zhonghua Wu, Taizhe Tan, Guosheng Lin et al.ACM MM 2021 · 27 citations
- Towards Multi-Pose Guided Virtual Try-On NetworkHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang et al.ICCV 2019 · 226 citations
