To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning
Ming Li, Jike Zhong, Shitian Zhao, Yuxiang Lai, Haoquan Zhang, Wang Bill Zhu, Kaipeng Zhang
摘要
This paper investigates the role of explicit thinking process in rule-based reinforcement fine-tuning (RFT) for multi-modal large language models (MLLMs). We first extend Thinking-RFT to image classification task, using verifiable rewards for fine-tuning (FT). Experiments show Thinking-RFT significantly outperforms supervised FT and yields a cross-dataset generalization effect. We then rethink and question whether explicit thinking in RFT is always necessary and beneficial. Challenging the convention that explicit thinking is crucial for the success of RFT, we introduce No-Thinking-RFT, exploring RFT without thinking by introducing a simple equality accuracy reward. We evaluate No-Thinking-RFT on six diverse tasks across different model sizes and types. Experiment results reveal four key findings: (1). Visual perception tasks do not require thinking during RFT, as No-Thinking-RFT consistently outperforms or matches Thinking-RFT across model sizes and types. (2). Models with limited capabilities struggle to generate highquality CoT for RFT, making Thinking-RFT less effective than No-Thinking-RFT.
(3). There are inconsistencies between the answers in the thinking tags and answer tags for some responses of Thinking-RFT, which show lower average accuracy than the overall accuracy. (4). The performance gain of No-Thinking-RFT mainly stems from improved learning during no thinking FT and the avoidance of inference overthinking, as evidenced by the partial gains from appending empty thinking tags at inference time of Thinking-RFT. We hypothesize that explicit thinking before verifiable answers may hinder reward convergence and reduce performance in certain scenarios. To test this, we propose Think-After-Answer, which places thinking after the answer to mitigate this effect for experimental verification. Lastly, we conduct a pilot study to explore whether MLLMs can learn when to think during RFT, introducing an Adaptive-Thinking method. Experiments show that model converges to either thinking or not depending on model capability, achieving comparable or better performance than both Thinking and No-Thinking-RFT. Our findings suggest MLLMs can adaptively decide to think or not based on their capabilities and task complexity, offering insights into the thinking process in RFT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- PhyCritic: Multimodal Critic Models for Physical AITianyi Xiong, Shihao Wang, Guilin Liu, Yi Dong 等CVPR 2026 · 被引用 12 次
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning GeneralizationYifan Du, Kun Zhou, Yingqian Min, Yue Ling 等CVPR 2026 · 被引用 7 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality AssessmentZiheng Jia, Jiaying Qian, Zicheng Zhang, Zijian Chen 等AAAI 2026 · 被引用 3 次
- UFT: Unifying Supervised and Reinforcement Fine-TuningMingyang Liu, Gabriele Farina, Asuman OzdaglarNeurIPS 2025 · 被引用 61 次
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 被引用 17 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- ProxyThinker: Test-Time Guidance through Small Visual ReasonersZilin Xiao, Jaywon Koo, Siru Ouyang, Jefferson Hernandez 等ICLR 2026 · 被引用 8 次
