Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
Jongwoo Ko, Sungnyun Kim, Sungwoo Cho, Se-Young Yun
摘要
Human-generated reward signals are critical for aligning generative models with human preferences, guiding both training and inference-time evaluations. While large language models (LLMs) employed as proxy evaluators, i.e., LLM-as-a-Judge, significantly reduce the costs associated with manual annotations, they typically require extensive modality-specific training data and fail to generalize well across diverse multimodal tasks. In this paper, we propose Flex-Judge, a reasoning-guided multimodal judge model that leverages minimal textual reasoning data to robustly generalize across multiple modalities and evaluation formats. Our core intuition is that structured textual reasoning explanations inherently encode generalizable decision-making patterns, enabling an effective transfer to multimodal judgments, e.g., with images or videos. Empirical results demonstrate that Flex-Judge, despite being trained on significantly fewer text data, achieves competitive or superior performance compared to state-of-the-art commercial APIs and extensively trained multimodal evaluators. Notably, Flex-Judge presents broad impact in modalities like molecule, where comprehensive evaluation benchmarks are scarce, underscoring its practical value in resource-constrained domains. Our framework highlights reasoning-based text supervision as a powerful, cost-effective alternative to traditional annotation-intensive approaches, substantially advancing scalable multimodal model-as-a-judge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual ReasoningQi Song, Honglin Li, Yingchen Yu, Haoyi Zhou 等CVPR 2026 · 被引用 16 次
- Stable and Efficient Single-Rollout RL for Multimodal ReasoningRui Liu, Dian Yu, Lei Ke, Haolin Liu 等CVPR 2026 · 被引用 13 次
- FairJudge : An Adaptive, Debiased, and Consistent LLM-as-a-JudgeBo Yang, Lanfei Feng, Yunkui Chen, Xiao Xu 等ICML 2026 · 被引用 3 次
它引用的顶会 Paper38
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- LLM-as-a-Judge for Reliable and Explainable Offline Evaluation in Top-K RecommendationYue Que, Junyi Zhou, Xiaokun Zhang, Haiming Jin 等KDD 2026
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video UnderstandingAbdul Waheed, Zhen Wu, Dareen Safar Alharthi, Seungone Kim 等ICLR 2026 · 被引用 4 次
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process JudgesJiaxin Ai, Pengfei Zhou, Zhaopan Xu, Ming Li 等ICCV 2025 · 被引用 9 次
- Improve LLM-as-a-Judge Ability as a General AbilityJiachen Yu, Shaoning Sun, Xiaohui Hu, Jiaxu Yan 等EMNLP 2025 · 被引用 1 次
- Think-J: Learning to Think for Generative LLM-as-a-JudgeHui Huang, Yancheng He, Hongli Zhou, Rui Zhang 等AAAI 2026 · 被引用 12 次
