R2-MultiOmnia: Leading Multilingual Multimodal Reasoning via Self-Training
Leonardo Ranaldi, Federico Ranaldi, Giulia Pucci
Abstract
Reasoning is an intricate process that tran-scends both language and vision; because of its inherently modality-agnostic nature, developing effective multilingual and multimodal reasoning capabilities is a substantial challenge for Multimodal Large Language Models (MLLMs). They struggle to activate complex reasoning behaviours, delivering step-wise explanation, questioning and reflection, particularly in multilingual settings where high-quality supervision across languages is lacking. Recent works have introduced eclectic strategies to enhance MLLMs’ reasoning; however, they remain related to a single language. To make MLLMs’ reasoning capabilities aligned among languages and improve modality performances, we propose R2-MultiOmnia , a modular approach that instructs the models to abstract key elements of the reasoning process and then refine reasoning trajectories via self-correction. Specifically, we instruct the models producing multimodal synthetic demonstrations by bridging modalities and then self-improving their capabilities. To stabilise learning and the reasoning processes structure, we propose Curriculum Learning Reasoning Stabilisation with structured output rewards to gradually refine the models’ capabilities to learn and deliver robust reasoning processes. Experiments show that R2-MultiOmnia improves multimodal reasoning, gets aligned performances among the languages approaching strong models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4558cb99-4bb6-468d-8c69-06478c82e6d1Cited by top-tier papers2
- Improving Multilingual Retrieval-Augmented Language Models through Dialectic Reasoning ArgumentationsLeonardo Ranaldi, Federico Ranaldi, Fabio Massimo Zanzotto, Barry Haddow et al.EMNLP 2025
- Advancing Oversight Reasoning across Languages for Audit Sycophantic Behaviour via X-AgentGiulia Pucci, Leonardo RanaldiEMNLP 2025
Builds on7
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree SearchHuanjin Yao, Jiaxing Huang, Wenhao Wu, Jingyi Zhang et al.NeurIPS 2025 · 147 citations
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li et al.ICCV 2025 · 37 citations
- R1-Onevision: Advancing Generalized Multimodal Reasoning Through Cross-Modal FormalizationYi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang et al.ICCV 2025 · 21 citations
Related papers
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image ReasoningHang Du, Jiayang Zhang, Guoshun Nan, Wendi Deng et al.ICCV 2025 · 1 citation
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual ReasoningYana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin et al.NeurIPS 2025 · 39 citations
- GThinker: Towards General Multimodal Reasoning via Cue-Guided RethinkingYufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue et al.CVPR 2026 · 14 citations
- Rationale-Enhanced Decoding for Multi-modal Chain-of-ThoughtShin'ya Yamaguchi, Kosuke Nishida, Daiki ChijiwaCVPR 2026
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
