What they do when in doubt: a study of inductive biases in seq2seq learners
Eugene Kharitonov, Rahma Chaabouni
Abstract
Sequence-to-sequence (seq2seq) learners are widely used, but we still have only limited knowledge about what inductive biases shape the way they generalize. We address that by investigating how popular seq2seq learners generalize in tasks that have high ambiguity in the training data. We use SCAN and three new tasks to study learners' preferences for memorization, arithmetic, hierarchical, and compositional reasoning. Further, we connect to Solomonoff's theory of induction and propose to use description length as a principled and sensitive measure of inductive biases. In our experimental study, we find that LSTM-based learners can learn to perform counting, addition, and multiplication by a constant from a single training example. Furthermore, Transformer and LSTM-based learners show a bias toward the hierarchical induction over the linear one, while CNN-based learners prefer the opposite. On the SCAN dataset, we find that CNN-based, and, to a lesser degree, Transformer- and LSTM-based learners have a preference for compositional generalization over memorization. Finally, across all our experiments, description length proved to be a sensitive measure of inductive biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa2a80e7-765f-4dca-90f3-232c90b7296cCited by top-tier papers7
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
- The King Is Naked: On the Notion of Robustness for Natural Language ProcessingEmanuele La Malfa, Marta KwiatkowskaAAAI 2022 · 31 citations
- Multiple Thinking Achieving Meta-Ability Decoupling for Object NavigationRonghao Dang, Lu Chen, Liuyi Wang, Zongtao He et al.ICML 2023 · 16 citations
- Self-attention Networks Localize When QK-eigenspectrum ConcentratesHan Bao, Ryuichiro Hataya, Ryo KarakidaICML 2024 · 16 citations
- A Survey of Inductive Reasoning for Large Language ModelsKedi Chen, Dezhao Ruan, Yuhao Dan, Yaoting Wang et al.ACL 2026 · 5 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Identity Crisis: Memorization and Generalization Under Extreme OverparameterizationChiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C. Mozer et al.ICLR 2020 · 96 citations
- A Formal Hierarchy of RNN ArchitecturesWilliam Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz et al.ACL 2020 · 6 citations
Related papers
- Mutual Exclusivity Training and Primitive Augmentation to Induce CompositionalityYichen Jiang, Xiang Zhou, Mohit BansalEMNLP 2022 · 1 citation
- Language Models Need Inductive Biases to Count InductivelyYingshan Chang, Yonatan BiskICLR 2025
- SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by SimulationMatthias Lindemann, Alexander Koller, Ivan TitovACL 2024 · 2 citations
- Information Locality as an Inductive Bias for Neural Language ModelsTaiga Someya, Anej Svete, Brian DuSell, Timothy J. O'Donnell et al.ACL 2025 · 6 citations
- Neural-Symbolic Recursive Machine for Systematic GeneralizationQing Li, Yixin Zhu, Yitao Liang, Ying Nian Wu et al.ICLR 2024 · 15 citations
