Understanding Deep Architecture with Reasoning Layer
Xinshi Chen, Yufei Zhang, Christoph Reisinger, Le Song
摘要
Recently, there has been a surge of interest in combining deep learning models with reasoning in order to handle more sophisticated learning tasks. In many cases, a reasoning task can be solved by an iterative algorithm. This algorithm is often unrolled, and used as a specialized layer in the deep architecture, which can be trained end-to-end with other neural components. Although such hybrid deep architectures have led to many empirical successes, the theoretical foundation of such architectures, especially the interplay between algorithm layers and other neural layers, remains largely unexplored. In this paper, we take an initial step towards an understanding of such hybrid deep architectures by showing that properties of the algorithm layers, such as convergence, stability, and sensitivity, are intimately related to the approximation and generalization abilities of the end-to-end model. Furthermore, our analysis matches closely our experimental observations under various conditions, suggesting that our theory can provide useful guidelines for designing deep architectures with reasoning layers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Implicit MLE: Backpropagating Through Discrete Exponential Family DistributionsMathias Niepert, Pasquale Minervini, Luca FranceschiNeurIPS 2021 · 被引用 121 次
- Neural Structured Prediction for Inductive Node ClassificationMeng Qu, Huiyu Cai, Jian TangICLR 2022 · 被引用 23 次
- Adaptive Perturbation-Based Gradient Estimation for Discrete Latent Variable ModelsPasquale Minervini, Luca Franceschi, Mathias NiepertAAAI 2023 · 被引用 16 次
- DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsYuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng 等NSDI 2026 · 被引用 14 次
- M-L2O: Towards Generalizable Learning-to-Optimize by Test-Time Fast Self-AdaptationJunjie Yang, Xuxi Chen, Tianlong Chen, Zhangyang Wang 等ICLR 2023
它引用的顶会 Paper5
- Generalization and Representational Limits of Graph Neural NetworksVikas K. Garg, Stefanie Jegelka, Tommi S. JaakkolaICML 2020 · 被引用 363 次
- Differentiation of Blackbox Combinatorial SolversMarin Vlastelica Pogancic, Anselm Paulus, Vít Musil, Georg Martius 等ICLR 2020 · 被引用 341 次
- MIPaaL: Mixed Integer Program as a LayerAaron M. Ferber, Bryan Wilder, Bistra Dilkina, Milind TambeAAAI 2020 · 被引用 169 次
- RNA Secondary Structure Prediction By Learning Unrolled AlgorithmsXinshi Chen, Yu Li, Ramzan Umarov, Xin Gao 等ICLR 2020 · 被引用 134 次
- GLAD: Learning Sparse Graph RecoveryHarsh Shrivastava, Xinshi Chen, Binghong Chen, Guanghui Lan 等ICLR 2020 · 被引用 39 次
相关 Paper
- Learning Iterative Reasoning through Energy MinimizationYilun Du, Shuang Li, Joshua B. Tenenbaum, Igor MordatchICML 2022 · 被引用 37 次
- From Growing to Looping: A Unified View of Iterative Computation in LLMsFerdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald 等ICML 2026 · 被引用 2 次
- Rethinking Deep Thinking: Stable Learning of Algorithms using Lipschitz ConstraintsJay Bear, Adam Prügel-Bennett, Jonathon S. HareNeurIPS 2024 · 被引用 10 次
- CombOptNet: Fit the Right NP-Hard Problem by Learning Integer Programming ConstraintsAnselm Paulus, Michal Rolínek, Vít Musil, Brandon Amos 等ICML 2021 · 被引用 73 次
- End-to-end Algorithm Synthesis with Recurrent Networks: Extrapolation without OverthinkingArpit Bansal, Avi Schwarzschild, Eitan Borgnia, Zeyad Emam 等NeurIPS 2022 · 被引用 54 次
