MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, Lichao Sun
Abstract
Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to the absence of multimodal benchmarks that align with human preferences. Drawing inspiration from the concept of LLM-as-a-Judge within LLMs, this paper introduces a novel benchmark, termed MLLM-as-a-Judge, to assess the ability of MLLMs in assisting judges across diverse modalities, encompassing three distinct tasks: Scoring Evaluation, Pair Comparison, and Batch Ranking. Our study reveals that, while MLLMs demonstrate remarkable human-like discernment in Pair Comparison, there is a significant divergence from human preferences in Scoring Evaluation and Batch Ranking. Furthermore, a closer examination reveals persistent challenges in the judgment capacities of LLMs, including diverse biases, hallucinatory responses, and inconsistencies in judgment, even in advanced models such as GPT-4V. These findings emphasize the pressing need for enhancements and further research efforts to be undertaken before regarding MLLMs as fully reliable evaluators. In light of this, we advocate for additional efforts dedicated to supporting the continuous development within the domain of MLLM functioning as judges. The code and dataset are publicly available at our project homepage: https://mllm-judge.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5280f49-9e57-46d8-914a-442dc9aff93bCited by top-tier papers107
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement LearningYifan Zhang, Xingyu Lu, Xiao Hu, Chaoyou Fu et al.ICLR 2026 · 65 citations
- Margin-Aware Preference Optimization for Aligning Diffusion Models Without ReferenceJiwoo Hong, Sayak Paul, Noah Lee, Kashif Rasul et al.AAAI 2026 · 43 citations
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi et al.NeurIPS 2025 · 39 citations
- WebCode2M: A Real-World Dataset for Code Generation from Webpage DesignsYi Gui, Zhen Li, Yao Wan, Yemin Shi et al.WWW 2025 · 38 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
Related papers
- MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-JudgeSua Lee, Sanghee Park, Jinbae ImACL 2026 · 1 citation
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video UnderstandingAbdul Waheed, Zhen Wu, Dareen Safar Alharthi, Seungone Kim et al.ICLR 2026 · 4 citations
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGIKaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li et al.ICML 2024 · 184 citations
- On Path to Multimodal Generalist: General-Level and General-BenchHao Fei, Yuan Zhou, Juncheng Li, Xiangtai Li et al.ICML 2025
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMsXuanwen Ding, Chengjun Pan, Zejun Li, Jiwen Zhang et al.ACL 2026 · 1 citation
