Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
Yao DU, Shanshan Song, Xiaomeng Li
Abstract
Multimodal large language models (MLLMs) struggle with numerical regression under longtailed target distributions. Token-level supervised fine-tuning (SFT) and point-wise regression rewards bias learning toward high-density regions, leading to regression-to-the-mean behavior and poor tail performance. We identify the lack of cross-sample relational supervision as a key limitation of existing MLLM training paradigms. To address it, we propose a distribution-aware reinforcement learning framework based on Group Relative Policy Optimization, which introduces batch-level comparison-based supervision via the Concordance Correlation Coefficient-based reward to align predicted and ground-truth distributions in terms of correlation, scale, and mean. The framework is plug-and-play, requiring no architectural modification. Experiments on a unified suite of long-tailed regression benchmarks show consistent improvements over SFT and existing MLLM regression methods, with particularly strong gains in medium- and few-shot regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Delving into Deep Imbalanced RegressionYuzhe Yang, Kaiwen Zha, Ying-Cong Chen, Hao Wang et al.ICML 2021 · 385 citations
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang et al.NeurIPS 2025 · 117 citations
- Perception-R1: Pioneering Perception Policy with Reinforcement LearningEn Yu, Kangheng Lin, Liang Zhao, Jisheng Yin et al.NeurIPS 2025 · 115 citations
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankTianhe Wu, Jian Zou, Jie Liang, Lei Zhang et al.NeurIPS 2025 · 92 citations
Related papers
- REAL: Regression-Aware Reinforcement Learning for LLM-as-a-JudgeYasi Zhang, Tianyu Chen, Mingyuan Zhou, Oscar Leong et al.ICML 2026
- DEVA: Fine-tuning Multimodal Large Language Models for Visual Perception TasksDebasmit Das, Munawar Hayat, Fatih PorikliCVPR 2026
- Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation ComprehensionLin Li, Wei Chen, Jiahui Li, Kwang-Ting Cheng et al.AAAI 2026 · 6 citations
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG et al.CVPR 2026 · 7 citations
