Revealing Procedural Reasoning Structures in Chain-of-Thought Training via Span-Level Gradient Organization
Jia Liu, Jiaxin Luo, Weiwen Xu, Jonathan M. Garibaldi, Xiaokun Wu, Yixue Hao, Min Chen
摘要
Chain-of-Thought (CoT) prompting enables large language models to produce multi-step reasoning, yet how such reasoning-related structure is expressed during training remains poorly understood. We present Gradient-based Structural Developer (GSD), an unsupervised framework with a principled gradient aggregation view that tracks span-level gradient during fine-tuning on reasoning benchmarks to understand how models develop structured, stepby-step reasoning capabilities. Our analysis shows that while gradients at the level of individual tokens are often noisy, aggregating gradients over contiguous reasoning-related spans reveals stable and recurring directional alignment across samples. We refer to these directionally aligned patterns as aligned sequential stresses, reflecting consistent gradient organization associated with similar reasoning procedures. Beyond capturing semantically similar reasoning instances, such gradient alignment also reveals structurally similar but semantically diverse cases that share common procedural organization. These findings position GSD as a diagnostic framework for analyzing how procedural reasoning structures emerge during training, with downstream selection results serving as auxiliary evidence correlating gradient alignment with adaptation efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMsXuan Zhang, Chao Du, Tianyu Pang, Qian Liu 等NeurIPS 2024 · 被引用 177 次
- Improving Chain-of-Thought Reasoning via Quasi-Symbolic AbstractionsLeonardo Ranaldi, Marco Valentino, André FreitasACL 2025 · 被引用 29 次
相关 Paper
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-ThoughtGuanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li 等NeurIPS 2025 · 被引用 8 次
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsXiang Zhang, Juntai Cao, Chenyu You, Dujian DingACL 2025 · 被引用 21 次
- Syntax vs. Semantics: How Transformers Learn Deep DependenciesJiangrui Zhao, Xiaoting DuICML 2026
- What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient PerspectiveMing Li, Yanhong Li, Tianyi ZhouACL 2025 · 被引用 27 次
- Exploring Layer Activation Dynamic of CoT via Knowledge ProbeChuanxin Zhang, Jiajun Liu, Yao He, Wenjun Ke 等ACL 2026
