Dynamic Weighted Combiner for Mixed-Modal Image Retrieval
Fuxiang Huang, Lei Zhang, Xiaowei Fu, Suqi Song
摘要
Mixed-Modal Image Retrieval (MMIR) as a flexible search paradigm has attracted wide attention. However, previous approaches always achieve limited performance, due to two critical factors are seriously overlooked. 1) The contribution of image and text modalities is different, but incorrectly treated equally. 2) There exist inherent labeling noises in describing users' intentions with text in web datasets from diverse real-world scenarios, giving rise to overfitting. We propose a Dynamic Weighted Combiner (DWC) to tackle the above challenges, which includes three merits. First, we propose an Editable Modality De-equalizer (EMD) by taking into account the contribution disparity between modalities, containing two modality feature editors and an adaptive weighted combiner. Second, to alleviate labeling noises and data bias, we propose a dynamic soft-similarity label generator (SSG) to implicitly improve noisy supervision. Finally, to bridge modality gaps and facilitate similarity learning, we propose a CLIP-based mutual enhancement module alternately trained by a mixed-modality contrastive loss. Extensive experiments verify that our proposed model significantly outperforms state-of-the-art methods on real-world datasets. The source code is available at https://github.com/fuxianghuang1/DWC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li 等AAAI 2026 · 被引用 12 次
- Aligning Composed Query with Image via Discriminative Perception from Negative CorrespondencesYifan Wang, Wuliang Huang, Chun YuanAAAI 2025 · 被引用 3 次
- TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text RetrievalShuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu 等AAAI 2025 · 被引用 2 次
- Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningHaomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng 等AAAI 2026 · 被引用 1 次
- Towards Robust Uncertainty Calibration for Composed Image RetrievalYifan Wang, Wuliang Huang, Yufan Wen, Shunning Liu 等NeurIPS 2025
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 被引用 344 次
- An Empirical Study of Training End-to-End Vision-and-Language TransformersZi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang 等CVPR 2022 · 被引用 313 次
- ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit SimilarityGinger Delmas, Rafael Sampaio de Rezende, Gabriela Csurka, Diane LarlusICLR 2022 · 被引用 147 次
相关 Paper
- Deep Evidential Learning with Noisy Correspondence for Cross-modal RetrievalYang Qin, Dezhong Peng, Xi Peng, Xu Wang 等ACM MM 2022 · 被引用 101 次
- Deep Evidential Hashing for Trustworthy Cross-Modal RetrievalYuan Li, Liangli Zhen, Yuan Sun, Dezhong Peng 等AAAI 2025 · 被引用 8 次
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei 等ACM MM 2023 · 被引用 53 次
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 被引用 12 次
- Image Search with Text Feedback by Deep Hierarchical Attention Mutual Information MaximizationChunbin Gu, Jiajun Bu, Zhen Zhang, Zhi Yu 等ACM MM 2021 · 被引用 33 次
