Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
Guijin Son, Jiwoo Hong, Hyunwoo Ko, James Thorne
摘要
Scaling pre-training compute has proven effective for achieving multilinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level problems in 55 languages. We test three test-time scaling methods-Outcome Reward Modeling (ORM), Process Reward Modeling (PRM), and Budget Forcing (BF)-on both Qwen2.5-1.5B Math and MR1-1.5B, a multilingual LLM we trained for extended reasoning. Our experiments show that using Qwen2.5-1.5B Math with ORM achieves a score of 35.8 on MCLM, while BF on MR1-1.5B attains 35.2. Although "thinking LLMs" have recently garnered significant attention, we find that their performance is comparable to traditional scaling methods like best-of-N once constrained to similar levels of inference FLOPs. Moreover, while BF yields a 20-point improvement on English AIME, it provides only a 1.94-point average gain across other languages-a pattern consistent across the other test-time scaling methods we studied-highlighting that test-time scaling may not generalize as effectively to multilingual tasks. To foster further research, we release MCLM, MR1-1.5B, and evaluation results. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- When AI Benchmarks Plateau: A Systematic Study of Benchmark SaturationMubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja 等ICML 2026 · 被引用 22 次
- Long Chain-of-Thought Reasoning Across LanguagesJosh Barua, Seun Eisape, Kayo Yin, Alane SuhrICLR 2026 · 被引用 21 次
- When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual ReasonersWeixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu 等NeurIPS 2025 · 被引用 20 次
- Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMsShulin Huang, Yiran Ding, Junshu Pan, Yue ZhangICLR 2026 · 被引用 11 次
- Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-ThoughtGuijin Son, Donghun Yang, Hitesh Laxmichand Patel, Amit Agarwal 等ICLR 2026 · 被引用 10 次
它引用的顶会 Paper22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- M-RewardBench: Evaluating Reward Models in Multilingual SettingsSrishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary 等ACL 2025
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li 等EMNLP 2025 · 被引用 33 次
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for ReasoningCharlie Victor Snell, Jaehoon Lee, Kelvin Xu, Aviral KumarICLR 2025
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu 等NeurIPS 2025 · 被引用 249 次
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui 等NeurIPS 2025 · 被引用 20 次
