DeepPCR: Parallelizing Sequential Operations in Neural Networks
Federico Danieli, Miguel Sarabia, Xavier Suau Cuadros, Pau Rodríguez, Luca Zappella
摘要
Parallelization techniques have become ubiquitous for accelerating inference and training of deep neural networks. Despite this, several operations are still performed in a sequential manner. For instance, the forward and backward passes are executed layer-by-layer, and the output of diffusion models is produced by applying a sequence of denoising steps. This sequential approach results in a computational cost proportional to the number of steps involved, presenting a potential bottleneck as the number of steps increases. In this work, we introduce DeepPCR, a novel algorithm which parallelizes typically sequential operations in order to speed up inference and training of neural networks. DeepPCR is based on interpreting a sequence of steps as the solution of a specific system of equations, which we recover using the Parallel Cyclic Reduction algorithm. This reduces the complexity of computing the sequential operations from to , thus yielding a speedup for large . To verify the theoretical lower complexity of the algorithm, and to identify regimes for speedup, we test the effectiveness of DeepPCR in parallelizing the forward and backward pass in multi-layer perceptrons, and reach speedups of up to for the forward and for the backward pass. We additionally showcase the flexibility of DeepPCR by parallelizing training of ResNets with as many as 1024 layers, and generation in diffusion models, enabling up to faster training and faster generation, respectively, when compared to the sequential approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language ModelsFederico Danieli, Pau Rodríguez, Miguel Sarabia, Xavier Suau 等ICLR 2026 · 被引用 18 次
- Predictability Enables Parallelization of Nonlinear State Space ModelsXavier Gonzalez, Leo Kozachkov, David M. Zoltowski, Kenneth L. Clarkson 等NeurIPS 2025 · 被引用 12 次
- Parallelizing MCMC Across the Sequence LengthDavid M. Zoltowski, Skyler Wu, Xavier Gonzalez, Leo Kozachkov 等NeurIPS 2025 · 被引用 6 次
- Accelerated training through iterative gradient propagation along the residual pathErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICLR 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
相关 Paper
- Accelerating Parallel Sampling of Diffusion ModelsZhiwei Tang, Jiasheng Tang, Hao Luo, Fan Wang 等ICML 2024 · 被引用 30 次
- Accelerating Diffusion Models via Parallel DenoisingYanming Chen, Zixin Ma, Chuanguang Yang, Zhulin An 等ACM MM 2025
- Parallel Sampling of Diffusion ModelsAndy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh 等NeurIPS 2023 · 被引用 144 次
- Communication-Efficient Diffusion Denoising Parallelization via Reuse-then-Predict MechanismKunyun Wang, Bohan Li, Kai Yu, Minyi Guo 等NeurIPS 2025 · 被引用 3 次
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 被引用 33 次
