Multimodal Compatibility Modeling via Exploring the Consistent and Complementary Correlations
Weili Guan, Haokun Wen, Xuemeng Song, Chung-Hsing Yeh, Xiaojun Chang, Liqiang Nie
Abstract
Existing methods towards outfit compatibility modeling seldom explicitly consider multimodal correlations. In this work, we explore the consistent and complementary correlations for better compatibility modeling. This is, however, non-trivial due to the following challenges: 1) how to separate and model these two kinds of correlations; 2) how to leverage the derived complementary cues to strengthen the text and vision-oriented representations of the given item; and 3) how to reinforce the compatibility modeling with text and vision-oriented representations. To address these challenges, we present a comprehensive multimodal outfit compatibility modeling scheme. It first nonlinearly projects each modality into separable consistent and complementary spaces via multi-layer perceptron, and then models the consistent and complementary correlations between two modalities by parallel and orthogonal regularization. Thereafter, we strengthen the visual and textual representation of items with complementary information, and further induct both the text-oriented and vision- oriented outfit compatibility modeling. We ultimately employ the mutual learning strategy to reinforce the final performance of compatibility modeling. Extensive experiments demonstrate the superiority of our scheme.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi et al.MobiCom 2022 · 94 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Personalized Fashion Compatibility Modeling via Metapath-guided Heterogeneous Graph LearningWeili Guan, Fangkai Jiao, Xuemeng Song, Haokun Wen et al.SIGIR 2022 · 51 citations
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 11 citations
- RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News DetectionXinquan Yu, Ziqi Sheng, Wei Lu, Xiangyang Luo et al.AAAI 2025 · 9 citations
Builds on5
- Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identificationYixiao Ge, Dapeng Chen, Hongsheng LiICLR 2020 · 651 citations
- Hierarchical Fashion Graph Network for Personalized Outfit RecommendationXingchen Li, Xiang Wang, Xiangnan He, Long Chen et al.SIGIR 2020 · 124 citations
- Learning Similarity Conditions Without Explicit SupervisionReuben Tan, Mariya I. Vasileva, Kate Saenko, Bryan A. PlummerICCV 2019 · 90 citations
- Comprehensive Linguistic-Visual Composition Network for Image RetrievalHaokun Wen, Xuemeng Song, Xin Yang, Yibing Zhan et al.SIGIR 2021 · 72 citations
- Fashion Compatibility Modeling through a Multi-modal Try-on-guided SchemeXue Dong, Jianlong Wu, Xuemeng Song, Hongjun Dai et al.SIGIR 2020 · 29 citations
Related papers
- Complementary Factorization towards Outfit Compatibility ModelingTianyu Su, Xuemeng Song, Na Zheng, Weili Guan et al.ACM MM 2021 · 18 citations
- Explainable Multi-Modality Alignment for Transferable RecommendationShenghao Yang, Weizhi Ma, Zhiqiang Guo, Min Zhang et al.WWW 2025 · 3 citations
- Text2Outfit: Controllable Outfit Generation With Multimodal Language ModelsYuanhao Zhai, Yen-Liang Lin, Minxu Peng, Larry S. Davis et al.ICCV 2025 · 1 citation
- Collocation and Try-on Network: Whether an Outfit is CompatibleNa Zheng, Xuemeng Song, Qingying Niu, Xue Dong et al.ACM MM 2021 · 23 citations
- Fashion Outfit Complementary Item RetrievalYen-Liang Lin, Son Dinh Tran, Larry S. DavisCVPR 2020
