The Emperor's New Reasoning: Format Imitation Overshadows Genuine Mathematical Understanding in SFT
Linyao Yang, Jian-Tao Huang, Yafei Lu, Zhenhui Jessie Li, Guirong Xue
摘要
Recent advances in large language models (LLMs) have yielded impressive gains on mathematical reasoning benchmarks via supervised fine-tuning (SFT). However, the brittleness of these models under input perturbations has cast doubt on whether such improvements reflect genuine reasoning abilities or merely superficial alignment with expected output formats. We investigate the mechanisms behind SFT improvements in small-scale LLMs, addressing four key questions: (1) Are performance gains primarily due to format alignment rather than reasoning? (2) Can high-quality supervision encourage genuine reasoning? (3) Does scaling data shift learning from format alignment to deeper reasoning? (4) Are format alignment gains consistent across model sizes and architectures? Through controlled experiments, we find that most performance improvements arise from format alignment rather than genuine reasoning enhancement. Moreover, SFT's effectiveness is strongly influenced by the alignment between the base model's inductive biases and the teacher model's output distribution, rather than the teacher's raw strength. Finally, scaling up training data offers diminishing returns and does not fundamentally alter the model's reasoning behavior. These findings suggest that current SFT practices may overestimate the reasoning abilities of LLMs and underscore the need for more rigorous evaluation methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- AlphaZero-Like Tree-Search can Guide Large Language Model Decoding and TrainingZiyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer 等ICML 2024 · 被引用 325 次
- Deductive Verification of Chain-of-Thought ReasoningZhan Ling, Yunhao Fang, Xuanlin Li, Zhiao Huang 等NeurIPS 2023 · 被引用 234 次
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
相关 Paper
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM ReasoningMaggie Ziyu Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu 等ICML 2026 · 被引用 102 次
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment QualityYuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki 等EMNLP 2025
- Self-Refine Instruction-Tuning for Aligning Reasoning in Language ModelsLeonardo Ranaldi, André FreitasEMNLP 2024 · 被引用 3 次
- Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use InsteadFeiyang Kang, Michael Kuchnik, Karthik Padthe, Marin Vlastelica 等ICLR 2026 · 被引用 27 次
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in ReasoningWang Yang, Zirui Liu, Hongye Jin, Qingyu Yin 等NeurIPS 2025 · 被引用 7 次
