Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models
Ethan Wilcox, Peng Qian, Richard Futrell, Ryosuke Kohita, Roger Levy, Miguel Ballesteros
Abstract
Humans can learn structural properties about a word from minimal experience, and deploy their learned syntactic representations uniformly in different grammatical contexts. We assess the ability of modern neural language models to reproduce this behavior in English and evaluate the effect of structural supervision on learning outcomes. First, we assess few-shot learning capabilities by developing controlled experiments that probe models' syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training. Second, we assess invariance properties of learned representation: the ability of a model to transfer syntactic generalizations from a base context (e.g., a simple declarative active-voice sentence) to a transformed context (e.g., an interrogative sentence). We test four models trained on the same dataset: an n-gram baseline, an LSTM, and two LSTM-variants trained with explicit structural supervision (Dyer et al., 2016; Charniak et al., 2016) . We find that in most cases, the neural models are able to induce the proper syntactic generalizations after minimal exposure, often from just two examples during training, and that the two structurally supervised models generalize more accurately than the LSTM model. All neural models are able to leverage information learned in base contexts to drive expectations in transformed contexts, indicating that they have learned some invariance properties of syntax.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98e610bf-c7b9-449c-b2fb-020e7d5cc25aCited by top-tier papers3
- miCSE: Mutual Information Contrastive Learning for Low-shot Sentence EmbeddingsTassilo Klein, Moin NabiACL 2023 · 10 citations
- Connecting degree and polarity: An artificial language learning studyLisa Bylinina, Alexey Tikhonov, Ekaterina GarmashEMNLP 2023
- Learning syntax without semantics: Disentangled tiny language modelsEzra Winston, Zico KolterICML 2026
Builds on1
Related papers
- Word Frequency Does Not Predict Grammatical Knowledge in Language ModelsCharles Yu, Ryan Sie, Nico Tedeschi, Leon BergenEMNLP 2020 · 6 citations
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 15 citations
- Structural Guidance for Transformer Language ModelsPeng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez AstudilloACL 2021
- Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNsKanishka Misra, Kyle MahowaldEMNLP 2024 · 11 citations
- Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic TransformationsMatthias Lindemann, Alexander Koller, Ivan TitovEMNLP 2024
