Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal Understanding
Xuan Lu, Kangle Li, Haohang Huang, Rui Meng, Wenjun Zeng, Xiaoyu Shen
Abstract
Recent advances in multimodal large language models (MLLMs) have substantially expanded the capabilities of multimodal retrieval, enabling systems to align and retrieve information across visual and textual modalities. Yet, existing benchmarks largely focus on coarse-grained or single-condition alignment, overlooking real-world scenarios where user queries specify multiple interdependent constraints across modalities. To bridge this gap, we introduce MCMR (Multi-Conditional Multimodal Retrieval): a largescale benchmark designed to evaluate fine-grained, multicondition cross-modal retrieval under natural-language queries. MCMR spans five product domains: upper and bottom clothing, jewelry, shoes, and furniture. It also preserves rich long-form metadata essential for compositional matching. Each query integrates complementary visual and textual attributes, requiring models to jointly satisfy all specified conditions for relevance. We benchmark a diverse suite of MLLM-based multimodal retrievers and vision-language rerankers to assess their condition-aware reasoning abilities. Experimental results reveal: (i) distinct modality asymmetries across models; (ii) visual cues dominate early-rank precision, while textual metadata stabilizes long-tail ordering; and (iii) MLLM-based pointwise rerankers markedly improve fine-grained matching by explicitly verifying query-candidate consistency. Overall, MCMR establishes a challenging and diagnostic benchmark for advancing multimodal retrieval toward compositional, constraint-aware, and interpretable understanding. Our code and dataset is available at https://github. com/EIT-NLP/MCMR
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91be643a-22ca-4e09-a8ba-4581c9525f6fBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He et al.ACL 2026 · 346 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
Related papers
- Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMSSheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin et al.ICLR 2025
- Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language ModelsXin Zhang, Yanzhao Zhang, Wen Xie, Mingxin Li et al.CVPR 2025
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal RetrievalSiyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao et al.ICLR 2026 · 13 citations
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah et al.CVPR 2025
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning ChainsXuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai et al.ICLR 2026 · 14 citations
