Leveraging pre-trained language models for linguistic analysis: A case of argument structure constructions
Hakyung Sung, Kristopher Kyle
Abstract
This study evaluates the effectiveness of pretrained language models in identifying argument structure constructions, important for modeling both first and second language learning. We examine three methodologies: (1) supervised training with RoBERTa using a gold-standard ASC treebank, including by-tag accuracy evaluation for sentences from both native and non-native English speakers, (2) prompt-guided annotation with and (3) generating training data through prompts with GPT-4, followed by RoBERTa training. Our findings indicate that RoBERTa trained on gold-standard data shows the best performance. While data generated through GPT-4 enhances training, it does not exceed the benchmarks set by gold-standard data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cdaebc4-e6ef-4774-a117-29e7071bdaa2Builds on1
Related papers
- Neural reality of argument structure constructionsBai Li, Zining Zhu, Guillaume Thomas, Frank Rudzicz et al.ACL 2022 · 38 citations
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 31 citations
- Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not ArgumentsMarc Feger, Katarina Boland, Stefan DietzeACL 2025
- Constructions are Revealed in Word DistributionsJoshua Rozner, Leonie Weissweiler, Kyle Mahowald, Cory ShainEMNLP 2025 · 8 citations
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 15 citations
