Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
Shijie Zhao, Xuanyu Zhang, Weiqi Li, Junlin Li, Li zhang, Tianfan Xue, Jian Zhang
Abstract
Reasoning-based image quality assessment (IQA) models trained through reinforcement learning (RL) exhibit exceptional generalization, yet the underlying mechanisms and critical factors driving this capability remain underexplored in current research. Moreover, despite their superior performance, these models incur inference energy usage and latency orders of magnitude higher than their earlier counterparts, restricting their deployment in specific scenarios. Through extensive experiments, this paper verifies and elaborates that through RL training, MLLMs leverage their reasoning capability to convert redundant visual representations into compact, cross-domain aligned text representations. This conversion is precisely the source of the generalization exhibited by these reasoning-based IQA models. Building on this fundamental insight, we propose a novel algorithm, RALI, which employs contrastive learning to directly align images with these generalizable text representations learned by RL. This approach eliminates the reliance on reasoning processes and even obviates the need to load an LLM during inference. For the quality scoring task, this framework achieves generalization performance comparable to reasoning-based models while requiring less than 5% of their model parameters and inference time. Code is available at RALI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a362b7d-70a1-4dce-960c-5d6598f108deCited by top-tier papers3
- UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and EnhancementWeiqi Li, Xuanyu Zhang, Bin Chen, Jingfen Xie et al.CVPR 2026 · 5 citations
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu et al.CVPR 2026 · 3 citations
- Return-Critic: Bridging Goal Discrepancy for Efficient Visual Reinforcement LearningRuyi Lu, Xuesong Wang, Hengrui Zhang, Yuhu ChengICML 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to RankTianhe Wu, Jian Zou, Jie Liang, Lei Zhang et al.NeurIPS 2025 · 92 citations
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal ReasoningYunhao Gou, Kai Chen, Zhili Liu, Lanqing HONG et al.ICLR 2026 · 7 citations
- Q-Insight: Understanding Image Quality via Visual Reinforcement LearningWeiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang et al.NeurIPS 2025 · 117 citations
- Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement LearningHaonan Jia, Shichao Dong, Xin Dong, Zenghui Sun et al.CVPR 2026
- UME-R1: Exploring Reasoning-Driven Generative Multimodal EmbeddingsZhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou et al.ICLR 2026 · 38 citations
