Understanding Deep Architecture with Reasoning Layer
Xinshi Chen, Yufei Zhang, Christoph Reisinger, Le Song
Abstract
Recently, there has been a surge of interest in combining deep learning models with reasoning in order to handle more sophisticated learning tasks. In many cases, a reasoning task can be solved by an iterative algorithm. This algorithm is often unrolled, and used as a specialized layer in the deep architecture, which can be trained end-to-end with other neural components. Although such hybrid deep architectures have led to many empirical successes, the theoretical foundation of such architectures, especially the interplay between algorithm layers and other neural layers, remains largely unexplored. In this paper, we take an initial step towards an understanding of such hybrid deep architectures by showing that properties of the algorithm layers, such as convergence, stability, and sensitivity, are intimately related to the approximation and generalization abilities of the end-to-end model. Furthermore, our analysis matches closely our experimental observations under various conditions, suggesting that our theory can provide useful guidelines for designing deep architectures with reasoning layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01278515-7942-4b9d-9421-965d9fd2ffe4Cited by top-tier papers6
- Implicit MLE: Backpropagating Through Discrete Exponential Family DistributionsMathias Niepert, Pasquale Minervini, Luca FranceschiNeurIPS 2021 · 121 citations
- Neural Structured Prediction for Inductive Node ClassificationMeng Qu, Huiyu Cai, Jian TangICLR 2022 · 23 citations
- Adaptive Perturbation-Based Gradient Estimation for Discrete Latent Variable ModelsPasquale Minervini, Luca Franceschi, Mathias NiepertAAAI 2023 · 16 citations
- DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsYuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng et al.NSDI 2026 · 14 citations
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang et al.ICLR 2023
Builds on5
- Generalization and Representational Limits of Graph Neural NetworksVikas K. Garg, Stefanie Jegelka, Tommi S. JaakkolaICML 2020 · 363 citations
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius et al.ICLR 2020 · 341 citations
- MIPaaL: Mixed Integer Program as a LayerAaron M. Ferber, Bryan Wilder, Bistra Dilkina, Milind TambeAAAI 2020 · 169 citations
- RNA Secondary Structure Prediction By Learning Unrolled AlgorithmsXinshi Chen, Yu Li, Ramzan Umarov, Xin Gao et al.ICLR 2020 · 134 citations
- GLAD: Learning Sparse Graph RecoveryHarsh Shrivastava, Xinshi Chen, Binghong Chen, Guanghui Lan et al.ICLR 2020 · 39 citations
Related papers
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 37 citations
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald et al.ICML 2026 · 2 citations
- Rethinking Deep Thinking: Stable Learning of Algorithms using Lipschitz ConstraintsJay Bear, Adam Prügel-Bennett, Jonathon S. HareNeurIPS 2024 · 10 citations
- CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming ConstraintsAnselm Paulus, Michal Rolínek, Vít Musil, Brandon Amos et al.ICML 2021 · 73 citations
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam et al.NeurIPS 2022 · 54 citations
