Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics
Sean Welleck, Peter West, Jize Cao, Yejin Choi
摘要
Neural sequence models trained with maximum likelihood estimation have led to breakthroughs in many tasks, where success is defined by the gap between training and test performance. However, their ability to achieve stronger forms of generalization remains unclear. We consider the problem of symbolic mathematical integration, as it requires generalizing systematically beyond the training set. We develop a methodology for evaluating generalization that takes advantage of the problem domain's structure and access to a verifier. Despite promising in-distribution performance of sequence-to-sequence models in this domain, we demonstrate challenges in achieving robustness, compositionality, and out-of-distribution generalization, through both carefully constructed manual test suites and a genetic algorithm that automatically finds large collections of failures in a controllable manner. Our investigation highlights the difficulty of generalizing well with the predominant modeling and learning approach, and the importance of evaluating beyond the test set, across different aspects of generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin 等ICLR 2024 · 被引用 189 次
- NaturalProver: Grounded Mathematical Proof Generation with Language ModelsSean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi 等NeurIPS 2022 · 被引用 108 次
- LILA: A Unified Benchmark for Mathematical ReasoningSwaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang 等EMNLP 2022 · 被引用 73 次
- SALSA: Attacking Lattice Cryptography with TransformersEmily Wenger, Mingjie Chen, François Charton, Kristin E. LauterNeurIPS 2022 · 被引用 61 次
它引用的顶会 Paper8
- Deep Learning For Symbolic MathematicsGuillaume Lample, François ChartonICLR 2020 · 被引用 477 次
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 被引用 149 次
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 被引用 121 次
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark 等ICLR 2020 · 被引用 109 次
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang 等EMNLP 2020 · 被引用 37 次
相关 Paper
- Structural generalization is hard for sequence-to-sequence modelsYuekun Yao, Alexander KollerEMNLP 2022 · 被引用 8 次
- ASyMOB: Algebraic Symbolic Mathematical Operations BenchmarkMichael Shalyt, Rotem Elimelech, Ido KaminerICML 2026 · 被引用 7 次
- Compositional Generalization via Neural-Symbolic Stack MachinesXinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song 等NeurIPS 2020 · 被引用 112 次
- Deep Generative Symbolic Regression with Monte-Carlo-Tree-SearchPierre-Alexandre Kamienny, Guillaume Lample, Sylvain Lamprier, Marco VirgolinICML 2023 · 被引用 50 次
- Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradientsBrenden K. Petersen, Mikel Landajuela, T. Nathan Mundhenk, Cláudio Prata Santiago 等ICLR 2021 · 被引用 444 次
