Lune

AAAI2026Top-tier venue

CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement Learning

Shun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie, Zhen Xu, Baoxun Wang

2026Year

Abstract

Compositional reasoning is a critical capability for multimodal models, enabling systematic understanding of complex scenes through structured combinations of objects, attributes, and relations. However, existing research on this ability primarily focuses on vision-language models (VLMs, e.g., CLIP and SigLIP), with limited exploration of multimodal large language models (MLLMs). To address this gap, we introduce CR 3 , a novel framework that enhances compositional reasoning abilities of MLLMs via rule-based reinforcement learning. CR 3 leverages rule-based rewards to optimize the MLLM's policy on systematically curated multimodal instruction-following tasks, guided by a modeladaptive dynamic task mixing strategy. Our approach boosts performance by over 19% on three compositional reasoning benchmarks, significantly outperforming supervised finetuning (SFT) by at least 12%. Crucially, CR 3 demonstrates superior generalization by improving performance on out-ofdomain benchmarks where SFT methods degrade, highlighting its effectiveness and data efficiency.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7bd55ae3-6aef-4a5f-ba0f-1f55801a47c4

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines