Lune

AAAI2026顶会

CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning

Shun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie, Zhen Xu, Baoxun Wang

2026年份

摘要

Compositional reasoning is a critical capability for multimodal models, enabling systematic understanding of complex scenes through structured combinations of objects, attributes, and relations. However, existing research on this ability primarily focuses on vision-language models (VLMs, e.g., CLIP and SigLIP), with limited exploration of multimodal large language models (MLLMs). To address this gap, we introduce CR 3 , a novel framework that enhances compositional reasoning abilities of MLLMs via rule-based reinforcement learning. CR 3 leverages rule-based rewards to optimize the MLLM's policy on systematically curated multimodal instruction-following tasks, guided by a modeladaptive dynamic task mixing strategy. Our approach boosts performance by over 19% on three compositional reasoning benchmarks, significantly outperforming supervised finetuning (SFT) by at least 12%. Crucially, CR 3 demonstrates superior generalization by improving performance on out-ofdomain benchmarks where SFT methods degrade, highlighting its effectiveness and data efficiency.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖