Multimodal Meta-Verifier with Explicit Structured Recalibration
Xinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi, Yizhen Zhang, Junhong Liu, Youliang Zhang, Zhiheng Li, Yujiu Yang, Ling Yang
摘要
Visual outcomes are increasingly central to multimodal large language models, making reliable and fine-grained verification essential for scaling generalist foundation models. In this work, we investigate multimodal meta-verification , which leverages verifier-generated rationales rather than decision-only signals, and explore how to effectively incorporate meta-verification feedback into multimodal verifier training. We identify two key findings. First, symbolic verifier outputs (e.g., bounding boxes) outperform textual explanations as meta-verification rationales, enabling efficient rule-based reinforcement learning rewards while avoiding reliance on model-based rewards from auxiliary judge models. Second, decoupling reinforcement learning objectives for binary judgment and meta-verification substantially outperforms joint reward optimization, due to intrinsic differences in output structure and learning dynamics. Based on these insights, we train OmniVerifier-M1 , a generalist visual verifier leveraging symbolic meta-verification and decoupled reinforcement learning. OmniVerifier-M1 provides robust verification and fine-grained error localization, and further enables M1-TTS , a verifier-driven agentic generation system achieving dynamic region-level self-correction. This approach paves the way for more reliable, interpretable, and fine-grained multimodal verification, supporting safer and more controllable foundation model deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language BenchmarkDongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang 等ICML 2024 · 被引用 345 次
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 等ICLR 2026 · 被引用 321 次
相关 Paper
- Visual Attention Reasoning via Hierarchical Search and Self-VerificationWei Cai, Jian Zhao, Yuchen Yuan, Tianle Zhang 等ACL 2026
- Generative Universal Verifier as Multimodal Meta-ReasonerXinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao 等ICLR 2026 · 被引用 20 次
- Bad Seeing or Bad Thinking? Rewarding Perception for Multimodal ReasoningHaozhe WANG, Qixin Xu, Changpeng Wang, Taofeng Xue 等ICML 2026 · 被引用 3 次
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao 等ACL 2026 · 被引用 4 次
- Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded VerificationMoises Andrade, Joonhyuk Cha, Brandon Ho, Vriksha Srihari 等ICLR 2026 · 被引用 12 次
