Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference Loss
Michael L. Iuzzolino, Michael C. Mozer, Samy Bengio
Abstract
Although deep feedforward neural networks share some characteristics with the primate visual system, a key distinction is their dynamics. Deep nets typically operate in serial stages wherein each layer completes its computation before processing begins in subsequent layers. In contrast, biological systems have cascaded dynamics: information propagates from neurons at all layers in parallel but transmission occurs gradually over time, leading to speed-accuracy trade offs even in feedforward architectures. We explore the consequences of biologically inspired parallel hardware by constructing cascaded ResNets in which each residual block has propagation delays but all blocks update in parallel in a stateful manner. Because information transmitted through skip connections avoids delays, the functional depth of the architecture increases over time, yielding anytime predictions that improve with internal-processing time. We introduce a temporal-difference training loss that achieves a strictly superior speed-accuracy profile over standard losses and enables the cascaded architecture to outperform state-of-the-art anytime-prediction methods. The cascaded architecture has intriguing properties, including: it classifies typical instances more rapidly than atypical instances; it is more robust to both persistent and transient noise than is a conventional ResNet; and its time-varying output trace provides a signal that can be exploited to improve information processing and inference. * Currently at Apple. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Handling Delay in Real-Time Reinforcement LearningIvan Anokhin, Rishav Rishav, Matthew Riemer, Stephen Chung et al.ICLR 2025 · 225 citations
- Adaptive recurrent vision performs zero-shot computation scaling to unseen difficulty levelsVijay Veerabadran, Srinivas Ravishankar, Yuan Tang, Ritik Raina et al.NeurIPS 2023 · 11 citations
Builds on5
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- Characterizing Structural Regularities of Labeled Data in Overparameterized ModelsZiheng Jiang, Chiyuan Zhang, Kunal Talwar, Michael C. MozerICML 2021 · 128 citations
- Learning To Stop While Learning To PredictXinshi Chen, Hanjun Dai, Yu Li, Xin Gao et al.ICML 2020 · 57 citations
- Resolution Adaptive Networks for Efficient InferenceLe Yang, Yizeng Han, Xi Chen, Shiji Song et al.CVPR 2020
Related papers
- Toward Practical Equilibrium Propagation: Brain-inspired Recurrent Neural Network with Feedback Regulation and Residual ConnectionsZhuo Liu, Tao ChenICLR 2026 · 3 citations
- Anytime Inference with Distilled Hierarchical Neural EnsemblesAdria Ruiz, Jakob VerbeekAAAI 2021 · 21 citations
- Latent Equilibrium: Arbitrarily fast computation with arbitrarily slow neuronsPaul Haider, Benjamin Ellenberger, Laura Kriener, Jakob Jordan et al.NeurIPS 2021 · 32 citations
- BioOSS: A Bio-Inspired Oscillatory State System with Spatio-Temporal DynamicsZhongju Yuan, Geraint A. Wiggins, Dick BotteldoorenNeurIPS 2025
- Towards Anytime Classification in Early-Exit Architectures by Enforcing Conditional MonotonicityMetod Jazbec, James Urquhart Allingham, Dan Zhang, Eric T. NalisnickNeurIPS 2023 · 21 citations
