Symbolic Brittleness in Sequence Models: On Systematic Generalization in Symbolic Mathematics
Sean Welleck, Peter West, Jize Cao, Yejin Choi
Abstract
Neural sequence models trained with maximum likelihood estimation have led to breakthroughs in many tasks, where success is defined by the gap between training and test performance. However, their ability to achieve stronger forms of generalization remains unclear. We consider the problem of symbolic mathematical integration, as it requires generalizing systematically beyond the training set. We develop a methodology for evaluating generalization that takes advantage of the problem domain's structure and access to a verifier. Despite promising in-distribution performance of sequence-to-sequence models in this domain, we demonstrate challenges in achieving robustness, compositionality, and out-of-distribution generalization, through both carefully constructed manual test suites and a genetic algorithm that automatically finds large collections of failures in a controllable manner. Our investigation highlights the difficulty of generalizing well with the predominant modeling and learning approach, and the importance of evaluating beyond the test set, across different aspects of generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fb01a01-4856-4500-bbba-0efb9f7fdc49Cited by top-tier papers10
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin et al.ICLR 2024 · 189 citations
- NaturalProver: Grounded Mathematical Proof Generation with Language ModelsSean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi et al.NeurIPS 2022 · 108 citations
- LILA: A Unified Benchmark for Mathematical ReasoningSwaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang et al.EMNLP 2022 · 73 citations
- SALSA: Attacking Lattice Cryptography with TransformersEmily Wenger, Mingjie Chen, François Charton, Kristin E. LauterNeurIPS 2022 · 61 citations
Builds on8
- Deep Learning For Symbolic MathematicsGuillaume Lample, François ChartonICLR 2020 · 477 citations
- COGS: A Compositional Generalization Challenge Based on Semantic InterpretationNajoung Kim, Tal LinzenEMNLP 2020 · 149 citations
- Selective Question Answering under Domain ShiftAmita Kamath, Robin Jia, Percy LiangACL 2020 · 121 citations
- Environmental drivers of systematicity and generalization in a situated agentFelix Hill, Andrew K. Lampinen, Rosalia Schneider, Stephen Clark et al.ICLR 2020 · 109 citations
- Consistency of a Recurrent Language Model With Respect to Incomplete DecodingSean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang et al.EMNLP 2020 · 37 citations
Related papers
- Structural generalization is hard for sequence-to-sequence modelsYuekun Yao, Alexander KollerEMNLP 2022 · 8 citations
- ASyMOB: Algebraic Symbolic Mathematical Operations BenchmarkMichael Shalyt, Rotem Elimelech, Ido KaminerICML 2026 · 7 citations
- Compositional Generalization via Neural-Symbolic Stack MachinesXinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song et al.NeurIPS 2020 · 112 citations
- Deep Generative Symbolic Regression with Monte-Carlo-Tree-SearchPierre-Alexandre Kamienny, Guillaume Lample, Sylvain Lamprier, Marco VirgolinICML 2023 · 50 citations
- Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradientsBrenden K. Petersen, Mikel Landajuela, T. Nathan Mundhenk, Cláudio Prata Santiago et al.ICLR 2021 · 444 citations
