Learning to Reason with Neural Networks: Generalization, Unseen Data and Boolean Measures
Emmanuel Abbe, Samy Bengio, Elisabetta Cornacchia, Jon M. Kleinberg, Aryo Lotfi, Maithra Raghu, Chiyuan Zhang
摘要
This paper considers the Pointer Value Retrieval (PVR) benchmark introduced in [ZRKB21], where a 'reasoning' function acts on a string of digits to produce the label. More generally, the paper considers the learning of logical functions with gradient descent (GD) on neural networks. It is first shown that in order to learn logical functions with gradient descent on symmetric neural networks, the generalization error can be lower-bounded in terms of the noise-stability of the target function, supporting a conjecture made in [ZRKB21]. It is then shown that in the distribution shift setting, when the data withholding corresponds to freezing a single feature (referred to as canonical holdout), the generalization error of gradient descent admits a tight characterization in terms of the Boolean influence for several relevant architectures. This is shown on linear models and supported experimentally on other models such as MLPs and Transformers. In particular, this puts forward the hypothesis that for such architectures and for learning logical functions such as PVR functions, GD tends to have an implicit bias towards low-degree representations, which in turn gives the Boolean influence for the generalization error under quadratic loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Generalization on the Unseen, Logic Reasoning and Degree CurriculumEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Kevin RizkICML 2023 · 被引用 68 次
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon 等NeurIPS 2024 · 被引用 52 次
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 被引用 15 次
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in TransformersYixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao 等NeurIPS 2025 · 被引用 10 次
- Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial StatesNoam Razin, Yotam Alexander, Edo Cohen-Karlik, Raja Giryes 等ICML 2024 · 被引用 5 次
它引用的顶会 Paper12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Accuracy on the Line: on the Strong Correlation Between Out-of-Distribution and In-Distribution GeneralizationJohn Miller, Rohan Taori, Aditi Raghunathan, Shiori Sagawa 等ICML 2021 · 被引用 323 次
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
相关 Paper
- An Initial Alignment between Neural Network and Target is Needed for Gradient Descent to LearnEmmanuel Abbe, Elisabetta Cornacchia, Jan Hazla, Christopher MarquisICML 2022 · 被引用 16 次
- Learning Language Representations with Logical Inductive BiasJianshu ChenICLR 2023
- Enhancing Transformers for Generalizable First-Order Logical EntailmentTianshi Zheng, Jiazheng Wang, Zihao Wang, Jiaxin Bai 等ACL 2025 · 被引用 7 次
- Injecting Logical Constraints into Neural Networks via Straight-Through EstimatorsZhun Yang, Joohyung Lee, Chiyoun ParkICML 2022 · 被引用 26 次
- A Functional Perspective on Learning Symmetric Functions with Neural NetworksAaron Zweig, Joan BrunaICML 2021 · 被引用 23 次
