Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
Satoki Ishikawa, Rio Yokota, Ryo Karakida
Abstract
Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters because of the locality, making it challenging to identify desirable settings in which the algorithm progresses in a stable manner. To provide theoretical and quantitative insights, we introduce the maximal update parameterization (µP) in the infinite-width limit for two representative designs of local targets: predictive coding (PC) and target propagation (TP). We verified that µP enables hyperparameter transfer across models of different widths. Furthermore, our analysis revealed unique and intriguing properties of µP that are not present in conventional BP. By analyzing deep linear networks, we found that PC's gradients interpolate between first-order and Gauss-Newton-like gradients, depending on the parameterization. We demonstrate that, in specific standard settings, PC in the infinite-width limit behaves more similarly to the first-order gradient. For TP, even with the standard scaling of the last layer, which differs from classical µP, its local loss optimization favors the feature learning regime over the kernel regime. Published as a conference paper at ICLR 2025 For standard BP, deep learning theory offers insights into the universal properties of learning (Bahri et al., 2020; Bartlett et al., 2021) . A key research focus in this area is understanding learning in the infinite-width limit, including studies on neural tangent kernel (NTK) and feature learning regimes (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 19 citations
- On the Infinite Width and Depth Limits of Predictive Coding NetworksFrancesco Innocenti, El Mehdi Achour, Rafal BogaczICML 2026
Builds on22
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- Can the Brain Do Backpropagation? - Exact Implementation of Backpropagation in Predictive Coding NetworksYuhang Song, Thomas Lukasiewicz, Zhenghua Xu, Rafal BogaczNeurIPS 2020 · 117 citations
- A Theoretical Framework for Target PropagationAlexander Meulemans, Francesco S. Carzaniga, Johan A. K. Suykens, João Sacramento et al.NeurIPS 2020 · 110 citations
- Towards Scaling Difference Target Propagation by Learning Backprop TargetsMaxence Ernoult, Fabrice Normandin, Abhinav Moudgil, Sean Spinney et al.ICML 2022 · 49 citations
Related papers
- CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement LearningAdam Lee, M Ganesh Kumar, Blake Bordelon, Cengiz PehlevanICML 2026
- Efficient Computation of Deep Nonlinear Infinite-Width Neural Networks that Learn FeaturesGreg Yang, Michael Santacroce, Edward J. HuICLR 2022 · 9 citations
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Reverse Differentiation via Predictive CodingTommaso Salvatori, Yuhang Song, Zhenghua Xu, Thomas Lukasiewicz et al.AAAI 2022 · 37 citations
- A Theoretical Framework for Inference and Learning in Predictive Coding NetworksBeren Millidge, Yuhang Song, Tommaso Salvatori, Thomas Lukasiewicz et al.ICLR 2023 · 8 citations
