An Analysis for Reasoning Bias of Language Models with Small Initialization
Junjie Yao, Zhongwang Zhang, Zhi-Qin John Xu
摘要
Transformer-based Large Language Models (LLMs) have revolutionized Natural Language Processing by demonstrating exceptional performance across diverse tasks. This study investigates the impact of the parameter initialization scale on the training behavior and task preferences of LLMs. We discover that smaller initialization scales encourage models to favor reasoning tasks, whereas larger initialization scales lead to a preference for memorization tasks. We validate this reasoning bias via real datasets and meticulously designed anchor functions. Further analysis of initial training dynamics suggests that specific model components, particularly the embedding space and self-attention mechanisms, play pivotal roles in shaping these learning biases. We provide a theoretical framework from the perspective of model training dynamics to explain these phenomena. Additionally, experiments on realworld language tasks corroborate our theoretical insights. This work enhances our understanding of how initialization strategies influence LLM performance on reasoning tasks and offers valuable guidelines for training models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- The Geometry of Reasoning: Flowing Logics in Representation SpaceYufa Zhou, Yixiao Wang, Xunjian Yin, Shuyan Zhou 等ICLR 2026 · 被引用 29 次
- From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training DynamicsZheng-An Chen, Tao LuoNeurIPS 2025 · 被引用 13 次
- Understanding LoRA as Knowledge Memory: An Empirical AnalysisSeungju Back, Dongwoo Lee, Naun Kang, Taehee Lee 等ICML 2026 · 被引用 10 次
- Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for ReasoningXinhao Yao, Ruifeng Ren, Yun Liao, Lizhong Ding 等ICLR 2026 · 被引用 6 次
- How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic InterpretabilityShawn Im, Changdae Oh, Zhen Fang, Sharon LiICLR 2026 · 被引用 4 次
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 被引用 181 次
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 被引用 113 次
相关 Paper
- Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or MemorizingZhongwang Zhang, Pengxiao Lin, Zhiwei Wang, Yaoyu Zhang 等NeurIPS 2024 · 被引用 20 次
- A Multi-Perspective Analysis of Memorization in Large Language ModelsBowen Chen, Namgi Han, Yusuke MiyaoEMNLP 2024 · 被引用 2 次
- Reasoning with Latent Thoughts: On the Power of Looped TransformersNikunj Saunshi, Nishanth Dikkala, Zhiyuan Li, Sanjiv Kumar 等ICLR 2025
- Neuron-Level Differentiation of Memorization and Generalization in Large Language ModelsKo-Wei Huang, Yi-Fu Fu, Ching-Yu Tsai, Yu-Chieh Tu 等EMNLP 2025
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
