Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
Satoki Ishikawa, Rio Yokota, Ryo Karakida
摘要
Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters because of the locality, making it challenging to identify desirable settings in which the algorithm progresses in a stable manner. To provide theoretical and quantitative insights, we introduce the maximal update parameterization (µP) in the infinite-width limit for two representative designs of local targets: predictive coding (PC) and target propagation (TP). We verified that µP enables hyperparameter transfer across models of different widths. Furthermore, our analysis revealed unique and intriguing properties of µP that are not present in conventional BP. By analyzing deep linear networks, we found that PC's gradients interpolate between first-order and Gauss-Newton-like gradients, depending on the parameterization. We demonstrate that, in specific standard settings, PC in the infinite-width limit behaves more similarly to the first-order gradient. For TP, even with the standard scaling of the last layer, which differs from classical µP, its local loss optimization favors the feature learning regime over the kernel regime. Published as a conference paper at ICLR 2025 For standard BP, deep learning theory offers insights into the universal properties of learning (Bahri et al., 2020; Bartlett et al., 2021) . A key research focus in this area is understanding learning in the infinite-width limit, including studies on neural tangent kernel (NTK) and feature learning regimes (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 被引用 19 次
- On the Infinite Width and Depth Limits of Predictive Coding NetworksFrancesco Innocenti, El Mehdi Achour, Rafal BogaczICML 2026
它引用的顶会 Paper22
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
- Can the Brain Do Backpropagation? - Exact Implementation of Backpropagation in Predictive Coding NetworksYuhang Song, Thomas Lukasiewicz, Zhenghua Xu, Rafal BogaczNeurIPS 2020 · 被引用 117 次
- A Theoretical Framework for Target PropagationAlexander Meulemans, Francesco S. Carzaniga, Johan A. K. Suykens, João Sacramento 等NeurIPS 2020 · 被引用 110 次
- Towards Scaling Difference Target Propagation by Learning Backprop TargetsMaxence Ernoult, Fabrice Normandin, Abhinav Moudgil, Sean Spinney 等ICML 2022 · 被引用 49 次
相关 Paper
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
- Efficient Computation of Deep Nonlinear Infinite-Width Neural Networks that Learn FeaturesGreg Yang, Michael Santacroce, Edward J. HuICLR 2022 · 被引用 9 次
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Reverse Differentiation via Predictive CodingTommaso Salvatori, Yuhang Song, Zhenghua Xu, Thomas Lukasiewicz 等AAAI 2022 · 被引用 37 次
- A Theoretical Framework for Inference and Learning in Predictive Coding NetworksBeren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz 等ICLR 2023 · 被引用 8 次
