UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
Tiancheng Gu, Kaicheng Yang, Kaichen Zhang, Xiang An, Ziyong Feng, Yueyi Zhang, Weidong Cai, Jiankang Deng, Lidong Bing
摘要
Universal multimodal embedding models are foundational to various tasks. Existing approaches typically employ inbatch negative mining by measuring the similarity of querycandidate pairs. However, these methods often struggle to capture subtle semantic differences among candidates and lack diversity in negative samples. Moreover, the embeddings exhibit limited discriminative ability in distinguishing false and hard negatives. In this paper, we leverage the advanced understanding capabilities of MLLMs to enhance representation learning and present a novel Universal Multimodal Embedding (UniME-V2) model. Our approach first constructs a potential hard negative set through global retrieval. We then introduce the MLLM-as-a-Judge mechanism, which utilizes MLLMs to assess the semantic alignment of query-candidate pairs and generate soft semantic matching scores. These scores serve as a foundation for hard negative mining, mitigating the impact of false negatives and enabling the identification of diverse, high-quality hard negatives. Furthermore, the semantic matching scores are used as soft labels to mitigate the rigid one-to-one mapping constraint. By aligning the similarity matrix with the soft semantic matching score matrix, the model learns semantic distinctions among candidates, significantly enhancing its discriminative capacity. To further improve performance, we propose UniME-V2-Reranker, a reranking model trained on our mined hard negatives through a joint pairwise and listwise optimization approach. We conduct comprehensive experiments on the MMEB benchmark and multiple retrieval tasks, demonstrating that our method achieves state-of-the-art performance on average across all tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ReMatch: Boosting Representation through Matching for Multimodal RetrievalQianying Liu, Xiao Liang, Zhiqiang Zhang, Yibo Chen 等CVPR 2026 · 被引用 8 次
- Very Efficient Listwise Multimodal Reranking for Long DocumentsYiqun Sun, Pengfei Wei, Lawrence HsiehICML 2026 · 被引用 1 次
- Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document RetrievalWeiqing Li, Jinyue Guo, Yaqi Wang, Haiyang Xiao 等CVPR 2026 · 被引用 1 次
- Illuminating Visual Identity in Universal Multimodal EmbeddingsJiawei Cao, Junyi Feng, Jiashen Hua, Ziheng Huang 等CVPR 2026 · 被引用 1 次
- Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image EditingTingyu Song, Yanzhao Zhang, Mingxin Li, Zhuoning Guo 等ACL 2026
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Unprecedented ScaleReza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li 等SC 2022 · 被引用 276 次
- Image Captioners Are Scalable Vision Learners TooMichael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai 等NeurIPS 2023 · 被引用 104 次
相关 Paper
- Mm-Embed: Universal Multimodal Retrieval with Multimodal LLMSSheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin 等ICLR 2025
- Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMsTiancheng Gu, Kaicheng Yang, Ziyong Feng, Xingjun Wang 等ACM MM 2025 · 被引用 6 次
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMsXiaojie Li, Chu Li, Shi-Zhe Chen, Xi ChenICLR 2026 · 被引用 10 次
- UME-R1: Exploring Reasoning-Driven Generative Multimodal EmbeddingsZhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou 等ICLR 2026 · 被引用 38 次
- SOLAR: Self-supervised Joint Learning for Symmetric Multimodal RetrievalWenjie Yang, Hang Yu, Yuyu Guo, Peng DiICML 2026
