UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
Xiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming Wu
摘要
The fashion domain includes a range of realworld multimodal tasks, such as multimodal retrieval and generation. Recent advancements in AI-generated content, particularly large language models for text and diffusion models for visuals, have spurred significant research interest in applying these multimodal models to fashion. However, fashion models must also effectively handle embedding tasks, like imageto-text and text-to-image retrieval. Moreover, current unified fashion models often lack the capability for image generation. In this work, we present UniFashion, a unified framework that tackles the challenges of multimodal generation and retrieval tasks in the fashion domain, by integrating image and text generation with retrieval tasks. UniFashion unifies embedding and generative processes through the use of a diffusion model and LLM, enabling controllable and high-fidelity generation. Our model significantly outperforms previous state-of-the-art models focused on single tasks across various fashion-related challenges and can be easily adapted to manage complex vision-language tasks. This study highlights the synergistic potential between multimodal generation and retrieval, offering a promising avenue for future research in the fashion domain. The source code is available at https: //github.com/xiangyu-mm/UniFashion .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAGDavid Anugraha, Patrick Amadeus Irawan, Anshul Singh, En-Shiun Annie Lee 等CVPR 2026 · 被引用 2 次
- WeatherGFM: Learning a Weather Generalist Foundation Model via In-context LearningXiangyu Zhao, Zhiwang Zhou, Wenlong Zhang, Yihao Liu 等ICLR 2025
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG SamplesQixin Sun, Ziqin Wang, Hengyuan Zhao, Yilin Li 等AAAI 2026
- MAI: A Multi-turn Aggregation-Iteration Model for Composed Image RetrievalYanzhe Chen, Zhiwen Yang, Jinglin Xu, Yuxin PengICLR 2025
- Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented GenerationJiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei 等EMNLP 2025
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- FashionDiff: A Controllable Diffusion Model Using Pairwise Fashion Elements for Intelligent DesignHan Yan, Haijun Zhang, Xiangyu Mu, Jicong Fan 等ACM MM 2023 · 被引用 17 次
- Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and BeyondTianxin Wei, Bowen Jin, Ruirui Li, Hansi Zeng 等ICLR 2024 · 被引用 46 次
- Unified Discrete Diffusion for Simultaneous Vision-Language GenerationMinghui Hu, Chuanxia Zheng, Zuopeng Yang, Tat-Jen Cham 等ICLR 2023 · 被引用 8 次
- Multimodal Garment Designer: Human-Centric Latent Diffusion Models for Fashion Image EditingAlberto Baldrati, Davide Morelli, Giuseppe Cartella, Marcella Cornia 等ICCV 2023 · 被引用 103 次
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha 等EMNLP 2022 · 被引用 9 次
