UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming Wu
Abstract
The fashion domain includes a range of realworld multimodal tasks, such as multimodal retrieval and generation. Recent advancements in AI-generated content, particularly large language models for text and diffusion models for visuals, have spurred significant research interest in applying these multimodal models to fashion. However, fashion models must also effectively handle embedding tasks, like imageto-text and text-to-image retrieval. Moreover, current unified fashion models often lack the capability for image generation. In this work, we present UniFashion, a unified framework that tackles the challenges of multimodal generation and retrieval tasks in the fashion domain, by integrating image and text generation with retrieval tasks. UniFashion unifies embedding and generative processes through the use of a diffusion model and LLM, enabling controllable and high-fidelity generation. Our model significantly outperforms previous state-of-the-art models focused on single tasks across various fashion-related challenges and can be easily adapted to manage complex vision-language tasks. This study highlights the synergistic potential between multimodal generation and retrieval, offering a promising avenue for future research in the fashion domain. The source code is available at https: //github.com/xiangyu-mm/UniFashion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cf0afd9-6a0d-4ce3-ba8c-c3fc1c6e822aCited by top-tier papers5
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee et al.CVPR 2026 · 2 citations
- WeatherGFM: Learning a Weather Generalist Foundation Model via In-context LearningXiangyu Zhao, Zhiwang Zhou, Wenlong Zhang, Yihao Liu et al.ICLR 2025
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG SamplesQixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li et al.AAAI 2026
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
- Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented GenerationJiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei et al.EMNLP 2025
Builds on36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- FashionDiff: A Controllable Diffusion Model Using Pairwise Fashion Elements for Intelligent DesignHan Yan, Haijun Zhang, Xiangyu Mu, Jicong Fan et al.ACM MM 2023 · 17 citations
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng et al.ICLR 2024 · 46 citations
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham et al.ICLR 2023 · 8 citations
- Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image EditingAlberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia et al.ICCV 2023 · 103 citations
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha et al.EMNLP 2022 · 9 citations
