POLIA: Policy Optimization with Visual-Object-Level Intrinsic Advantage for Multimodal Reasoning
Yiran Zeng, Da Chen, Hangyu Mao, Yuanxing Zhang, Pengfei Wan, Mengchen Zhao
摘要
Recent advances in group-based reinforcement learning (RL) greatly improve LLMs' ability in text reasoning. Yet, these methods lack sufficient modeling of multimodal information, leading to significant reasoning hallucination. In this work, we propose POLIA, a novel group-based RL method with visual-object-level intrinsic advantage for multimodal reasoning. POLIA introduces two advantage computation stages over candidate answers and visual objects, respectively. The answer-level extrinsic advantages are computed based on the extrinsic rewards of a group of candidate answers. Moreover, we compute an intrinsic advantage for each visual object based on its confidence score and reference relations with final answers. Intuitively, the intrinsic advantage of an object reflects its potential contribution to the correct answer. This two-stage advantage computation ensures an accurate credit assignment mechanism over multimodal reasoning sequences with multiple visual objects. Experimental results on diverse multimodal reasoning benchmarks show that POLIA significantly outperforms open MLLMs and strong baselines. Code is available at https://github.com/dudu115/POLIAcode.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Grounding Multimodal Large Language Models to the WorldZhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 等ICLR 2024 · 被引用 1,170 次
相关 Paper
- EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel GroundingYuejiao Su, Xinshen ZHANG, Zhen Ye, Lei Yao 等ICML 2026
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang 等CVPR 2026 · 被引用 3 次
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking ReasoningJunhao Shen, Haiteng Zhao, Yuzhe Gu, Songyang Gao 等NeurIPS 2025 · 被引用 13 次
- Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent GamesYidong He, Yutao Lai, Pengxu Yang, Jiarui Gan 等ICML 2026
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang 等NeurIPS 2025 · 被引用 3 次
