R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
Qi Yang, Bolin Ni, Shiming Xiang, Houwen Peng
摘要
Multimodal Large Language Models (MLLMs) with explicit step-by-step reasoning have achieved strong performance on complex tasks. However, such reasoning is unnecessary for many simple queries and introduces substantial computational overhead. To address this inefficiency, we present R-4B, an auto-thinking MLLM that dynamically determines whether to invoke the reasoning process based on input complexity.Our key idea is to equip a single model with both thinking and non-thinking capabilities and train it to select the appropriate mode. We first introduce bi-mode annealing, a unified training paradigm that constructs a model competent in both reasoning-intensive and direct-answer settings without requiring explicit complexity annotations. Building on this foundation, we propose Bi-mode Policy Optimization (BPO), a lightweight reinforcement learning algorithm that employs a dual-rollout mechanism: for each input, the model generates both thinking and non-thinking responses. This prevents mode collapse and enables robust learning of an adaptive reasoning policy using only simple, rule-based rewards.Extensive experiments across 25 benchmarks show that R-4B achieves state-of-the-art performance among models of similar scale. It consistently surpasses Qwen2.5-VL-7B and matches or exceeds larger models such as Kimi-VL-A3B-Thinking-2506 (16B) on reasoning-intensive tasks, while reducing computational cost by avoiding redundant reasoning. Our results demonstrate that adaptive auto-thinking offers an effective and scalable pathway toward more efficient multimodal reasoning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMsYi Zhang, Bolin Ni, Xin-Sheng Chen, Hengrui Zhang 等ICLR 2026 · 被引用 30 次
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering TwiceShuming Liu, Mingchen Zhuge, Changsheng Zhao, Jun Chen 等CVPR 2026 · 被引用 18 次
- OralGPT-Omni: A Versatile Dental Multimodal Large Language ModelJing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan 等CVPR 2026 · 被引用 11 次
- Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token InteractionsLin Chen, zhaoxiaoke, Kun Ding, Weiwei Feng 等ICML 2026 · 被引用 4 次
- AFM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid ReasoningQianben Chen, Jingyi Cao, Jiayu Zhang, Tianrui Qin 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper14
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang 等ICML 2024 · 被引用 1,191 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang 等ACL 2025 · 被引用 377 次
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang 等CVPR 2024 · 被引用 213 次
相关 Paper
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian 等NeurIPS 2025 · 被引用 69 次
- Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM ReasoningQinghe Ma, Zhen Zhao, Yiming Wu, Jian Zhang 等ICML 2026
- MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy OptimizationXiangyu Zhao, Junming Lin, Tianhao Liang, Yifan Zhou 等ICLR 2026 · 被引用 4 次
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong 等NeurIPS 2025 · 被引用 71 次
- Thinkless: LLM Learns When to ThinkGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2025 · 被引用 128 次
