UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
Wei Zhang, Yeying Jin, Xin Li, Yan Zhang, Xiaofeng Cong, Cong Wang, Fengcai Qiao, Zhichao Lian
摘要
Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance. The source code and pretrained models are available at https://github.com/zwplus/UniFit .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Any2anytryon: Leveraging Adaptive Position Embeddings for Versatile Virtual Clothing TasksHailong Guo, Bohan Zeng, Yiren Song, Wentao Zhang 等ICCV 2025 · 被引用 13 次
- VTON-VLLM: Aligning Virtual Try-On Models with Human PreferencesSiqi Wan, Jingwen Chen, Qi Cai, Yingwei Pan 等NeurIPS 2025 · 被引用 4 次
- OmniTry: Virtual Try-On Anything without MasksYutong Feng, Linlin Zhang, Hengyuan Cao, Yiming Chen 等NeurIPS 2025 · 被引用 16 次
- OmniVTON: Training-Free Universal Virtual Try-OnZhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li 等ICCV 2025 · 被引用 7 次
- Towards Multi-Pose Guided Virtual Try-On NetworkHaoye Dong, Xiaodan Liang, Xiaohui Shen, Bochao Wang 等ICCV 2019 · 被引用 226 次
