IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter
Xiaojing Zhong, Zhonghua Wu, Xiaofeng Yang, Guosheng Lin, Qingyao Wu
Abstract
Given a pair of images depicting a person and a garment separately, image-based 3D virtual try-on methods aim to reconstruct a 3D human model that realistically portrays the person wearing the desired garment. In this paper, we present IPV-TON, a novel image-based 3D virtual try-on framework. IPV-TON employs score distillation sampling with image prompts to optimize a hybrid 3D human representation, integrating target garment features into diffusion priors through an image prompt adapter. To avoid interference with non-target areas, we leverage mask-guided image prompt embeddings to focus the image features on the try-on regions. Moreover, we impose geometric constraints on the 3D model with a pseudo silhouette generated by ControlNet, ensuring that the clothed 3D human model retains the shape of the source identity while accurately wearing the target garments. Extensive qualitative and quantitative experiments demonstrate that IPVTON outperforms previous methods in image-based 3D virtual try-on tasks, excelling in both geometry and texture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bd68705-239e-413d-850b-a3140454eb02Cited by top-tier papers1
Ask how each one uses itBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content CreationRui Chen, Yongwei Chen, Ningxin Jiao, Kui JiaICCV 2023 · 769 citations
- Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape SynthesisTianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu et al.NeurIPS 2021 · 652 citations
Related papers
- DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion ModelsZhenyu Xie, Haoye Dong, Yufei Gao, Zehua Ma et al.ACM MM 2024 · 9 citations
- OOTDiffusion: Outfitting Fusion Based Latent Diffusion for Controllable Virtual Try-OnYuhao Xu, Tao Gu, Weifeng Chen, Arlene ChenAAAI 2025 · 177 citations
- Shape-Guided Clothing Warping for Virtual Try-OnXiaoyu Han, Shunyuan Zheng, Zonglin Li, Chenyang Wang et al.ACM MM 2024 · 5 citations
- Disentangled Cycle Consistency for Highly-Realistic Virtual Try-OnChongjian Ge, Yibing Song, Yuying Ge, Han Yang et al.CVPR 2021
- Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-OnXu Yang, Changxing Ding, Zhibin Hong, Junhao Huang et al.CVPR 2024 · 25 citations
