Lune

EMNLP2025Top-tier venue

Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational Knowledge

Tsedeniya Kinfe Temesgen, Marion Di Marco, Alexander Fraser

2025Year
2Citations

Abstract

This paper presents a study of the linguistic knowledge and generalization capabilities of Large Language Models (LLMs), focusing on their morphosyntactic competence. We design three diagnostic tasks: (i) labeling syntactic information at the sentence level -identifying subjects, objects, and indirect objects; (ii) derivational decomposition at the word level -identifying morpheme boundaries and labeling the decomposed sequence; and (iii) in-depth study of morphological decomposition in German and Amharic. We evaluate prompting strategies in GPT-4o and LLaMA 3.3-70B to extract different types of linguistic structures for typologically diverse languages. Our results show that GPT-4o consistently outperforms LLaMA in all tasks; however, both models exhibit limitations and show little evidence of abstract morphological rule learning. Importantly, we show strong evidence that the models fail to learn underlying morphological structures. Therefore, raising important doubts about their ability to generalize.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9568e4f8-fa95-4d2e-9daf-bcdcfed1ce99

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines