FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
Xiao Han, Xiatian Zhu, Licheng Yu, Li Zhang, Yi-Zhe Song, Tao Xiang
Abstract
In the fashion domain, there exists a variety of visionand-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in each individual input/output format and dataset size. It has been common to design a task-specific model and fine-tune it independently from a pre-trained V+L model (e.g., CLIP). This results in parameter inefficiency and inability to exploit inter-task relatedness. To address such issues, we propose a novel FAshion-focused Multi-task Efficient learning method for Vision-and-Language tasks (FAME-ViL) in this work. Compared with existing approaches, FAME-ViL applies a single model for multiple heterogeneous fashion tasks, therefore being much more parameter-efficient. It is enabled by two novel components: (1) a task-versatile architecture with cross-attention adapters and task-specific adapters integrated into a unified V+L model, and (2) a stable and effective multi-task training strategy that supports learning from heterogeneous data and prevents negative transfer. Extensive experiments on four fashion tasks show that our FAME-ViL can save 61.5% of parameters over alternatives, while significantly outperforming the conventional independently trained single-task models. Code is
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3f73ead-90d7-496e-b8dd-9af49ed5f9cfCited by top-tier papers22
- Sentence-level Prompts Benefit Composed Image RetrievalYang Bai, Xinxing Xu, Yong Liu, Salman Khan et al.ICLR 2024 · 75 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Controllable Person Image Synthesis with Pose-Constrained Latent DiffusionXiao Han, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song et al.ICCV 2023 · 36 citations
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei et al.SIGIR 2024 · 30 citations
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu et al.SIGIR 2024 · 29 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- FashionSAP: Symbols and Attributes Prompt for Fine-Grained Fashion Vision-Language Pre-TrainingYunpeng Han, Lisai Zhang, Qingcai Chen, Zhijian Chen et al.CVPR 2023
- FaD-VLP: Fashion Vision-and-Language Pre-training towards Unified Retrieval and CaptioningSuvir Mirchandani, Licheng Yu, Mengjiao Wang, Animesh Sinha et al.EMNLP 2022 · 9 citations
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsSanjoy Chowdhury, Sayan Nag, Dinesh ManochaEMNLP 2023 · 17 citations
- MmAP: Multi-Modal Alignment Prompt for Cross-Domain Multi-Task LearningYi Xin, Junlong Du, Qiang Wang, Ke Yan et al.AAAI 2024 · 102 citations
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin et al.AAAI 2024 · 94 citations
