ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
Yiran Guan, Sifan Tu, Dingkang Liang, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Jian Luan, Yuliang Liu, Xiang Bai
Abstract
Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing Omni-modal Large Language Models (OLLM) excel at perceiving diverse modalities, they lack the complex reasoning abilities of recent Large Reasoning Models (LRM). However, enhancing the reasoning ability of OLLMs through additional training presents significant challenges, including the need for high-quality data, task-specific adaptation, and substantial computational costs. To address these limitations, we propose ThinkOmni, a training-free framework that lifts textual reasoning to omni-modal scenarios. ThinkOmni introduces two key components: 1) LRM-as-a-Guide, which leverages off-the-shelf LRMs to guide the OLLM decoding process; 2) Stepwise Contrastive Scaling, which adaptively balances perception and reasoning signals without manual hyperparameter tuning. Experiments on six multimodality reasoning benchmarks demonstrate that ThinkOmni consistently delivers performance improvements, with main results achieving 70.2% on MathVista and 75.5% on MMAU. Overall, ThinkOmni offers a flexible and generalizable solution for omni-modal reasoning and provides new insights into the generalization and application of reasoning capabilities. Code is publicly available at https://github.com/1ranGuan/thinkomni Actually, this is not a trivial problem, and despite considerable efforts, existing approaches to omnimodal reasoning are still limited in several critical aspects. Specifically, 1) Insufficient modality diversity. Current studies largely focus on specific modalities (e.g., image (Liu et al., 2025b;a; Lin et al., 2025 ), audio (Li et al., 2025a), or video (Wang et al., 2025)), rather than generalizing across Work done at Xiaomi Inc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9288877-ba9b-4a02-8d65-ae8d02dfd3e7Cited by top-tier papers1
Ask how each one uses itBuilds on14
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionChaoyou Fu, Haojia Lin, Xiong Wang, Yifan Zhang et al.NeurIPS 2025 · 234 citations
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du et al.NeurIPS 2025 · 143 citations
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsWenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo et al.NeurIPS 2025 · 103 citations
Related papers
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPOYicheng Xiao, Lin Song, Yukang Chen, Yingmin Luo et al.NeurIPS 2025 · 34 citations
- GThinker: Towards General Multimodal Reasoning via Cue-Guided RethinkingYufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue et al.CVPR 2026 · 14 citations
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu et al.ICLR 2026 · 4 citations
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal ReasoningYucheng Wang, Yifan Hou, Aydin Javadov, Mubashara Akhtar et al.ICLR 2026 · 3 citations
- MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language ModelsSangyun Chung, Se Yeon Kim, Youngchae Chee, Yong Man RoCVPR 2026 · 5 citations
