Understanding Subword Compositionality of Large Language Models
Qiwei Peng, Yekun Chai, Anders Søgaard
Abstract
Large language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations. In this paper, we present a comprehensive set of experiments to probe how LLMs compose subword information, focusing on three key aspects: structural similarity, semantic decomposability, and form retention. Our analysis of the experiments suggests that these five LLM families can be classified into three distinct groups, likely reflecting difference in their underlying composition strategies. Specifically, we observe (i) three distinct patterns in the evolution of structural similarity between subword compositions and whole-word representations across layers; (ii) great performance when probing layer by layer their sensitivity to semantic decompositionality; and (iii) three distinct patterns when probing sensitivity to formal features, e.g., character sequence length. These findings provide valuable insights into the compositional dynamics of LLMs and highlight different compositional pattens in how LLMs encode and integrate subword information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 2 citations
- Autoregressive Pre-Training on Pixels and TextsYekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang et al.EMNLP 2024 · 1 citation
Related papers
- Are representations built from the ground up? An empirical examination of local composition in language modelsEmmy Liu, Graham NeubigEMNLP 2022 · 5 citations
- From Tokens to Words: On the Inner Lexicon of LLMsGuy Kaplan, Matanel Oren, Yuval Reif, Roy SchwartzICLR 2025
- Differential syntactic and semantic encoding in LLMsSantiago Acevedo, Alessandro Laio, Marco BaroniICML 2026 · 7 citations
- Language Models Learn Universal Representations of Numbers and Here's Why You Should CareMichal Stefánik, Timothee Mickus, Marek Kadlcík, Bertram Højer et al.ACL 2026 · 1 citation
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 2 citations
