Thyme: Think Beyond Images
Yifan Zhang, Xingyu Lu, Shukang Yin, Chaoyou Fu, Wei Chen, Xiao Hu, Bin Wen, Kaiyu Jiang, Changyi Liu, Tianke Zhang, Haonan Fan, Kaibing Chen
Abstract
Following OpenAI's introduction of the "thinking with images" concept, recent efforts have explored stimulating the use of visual information in the reasoning process to enhance model performance in perception and reasoning tasks. However, to the best of our knowledge, no open-source work currently offers a feature set as rich as proprietary models (OpenAI O3 (OpenAI, 2025)), which can perform diverse image manipulations and simultaneously enhance logical reasoning capabilities through code. In this paper, we make a preliminary attempt in this direction by introducing Thyme (Think Beyond Images), a novel paradigm for enabling multimodal large language models to transcend existing "think with images" approaches by autonomously generating and executing diverse image processing and computational operations via executable code (Figure 2 ). This approach not only facilitates a rich, on-the-fly set of image manipulations (e.g., cropping, rotation, contrast enhancement), but also allows for mathematical computations, all while maintaining high autonomy in deciding when and how to apply these operations. We activate this capability through a two-stage training strategy: an initial Supervised Fine-Tuning (SFT) on a curated dataset of 500K samples to teach code generation, followed by a Reinforcement Learning (RL) phase to refine decision-making. For the RL stage, we manually collect and design high-resolution question-answer pairs to increase the learning difficulty, and we propose GRPO-ATS (Group Relative Policy Optimization with Adaptive Temperature Sampling), an algorithm that applies distinct temperatures to text and code generation to balance reasoning exploration with code execution precision. We conduct extensive experimental analysis and ablation studies. As shown in Figure 1 , comprehensive evaluations on nearly 20 benchmarks show that Thyme yields significant and consistent performance gains, particularly in challenging high-resolution perception and complex reasoning tasks. We release our datasets, sandbox, and code to facilitate future research. Figure 1 : Benchmark performance of Thyme. The comprehensive set of image manipulation capabilities enables Thyme to achieve significant improvements over the baseline in perception tasks. By leveraging its ability to convert complex mathematical reasoning into executable code, it consistently outperforms baselines in mathematical reasoning benchmarks. Furthermore, the observed gains across a wide range of general benchmarks further validate the effectiveness of our training approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddd7b0b7-e34e-4f52-8349-dc055a515e18Cited by top-tier papers24
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang et al.CVPR 2026 · 35 citations
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationYunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin et al.CVPR 2026 · 31 citations
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy OptimizationXinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li et al.CVPR 2026 · 25 citations
- Skyra: AI-Generated Video Detection via Grounded Artifact ReasoningYifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun et al.CVPR 2026 · 24 citations
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal ReasoningChi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu et al.CVPR 2026 · 22 citations
Builds on20
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
Related papers
- Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language ModelsZhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng et al.ICML 2026 · 1 citation
- GThinker: Towards General Multimodal Reasoning via Cue-Guided RethinkingYufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue et al.CVPR 2026 · 14 citations
- Fact : Teaching MLLMs with Faithful, Concise and Transferable RationalesMinghe Gao, Shuang Chen, Liang Pang, Yuan Yao et al.ACM MM 2024 · 2 citations
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual ReasoningYana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin et al.NeurIPS 2025 · 39 citations
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual DrawingJunfei Wu, Jian Guan, Kaituo Feng, Qiang Liu et al.NeurIPS 2025 · 153 citations
