Understanding Subword Compositionality of Large Language Models
Qiwei Peng, Yekun Chai, Anders Søgaard
摘要
Large language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations. In this paper, we present a comprehensive set of experiments to probe how LLMs compose subword information, focusing on three key aspects: structural similarity, semantic decomposability, and form retention. Our analysis of the experiments suggests that these five LLM families can be classified into three distinct groups, likely reflecting difference in their underlying composition strategies. Specifically, we observe (i) three distinct patterns in the evolution of structural similarity between subword compositions and whole-word representations across layers; (ii) great performance when probing layer by layer their sensitivity to semantic decompositionality; and (iii) three distinct patterns when probing sensitivity to formal features, e.g., character sequence length. These findings provide valuable insights into the compositional dynamics of LLMs and highlight different compositional pattens in how LLMs encode and integrate subword information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 被引用 213 次
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky 等ICLR 2023 · 被引用 17 次
- Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding WordsJosef Klafka, Allyson EttingerACL 2020 · 被引用 2 次
- Autoregressive Pre-Training on Pixels and TextsYekun Chai, Qingyi Liu, Jingwu Xiao, Shuohuan Wang 等EMNLP 2024 · 被引用 1 次
相关 Paper
- Are representations built from the ground up? An empirical examination of local composition in language modelsEmmy Liu, Graham NeubigEMNLP 2022 · 被引用 5 次
- From Tokens to Words: On the Inner Lexicon of LLMsGuy Kaplan, Matanel Oren, Yuval Reif, Roy SchwartzICLR 2025
- Differential syntactic and semantic encoding in LLMsSantiago Acevedo, Alessandro Laio, Marco BaroniICML 2026 · 被引用 7 次
- Language Models Learn Universal Representations of Numbers and Here's Why You Should CareMichal Stefánik, Timothee Mickus, Marek Kadlcík, Bertram Højer 等ACL 2026 · 被引用 1 次
- Extracting Linguistic Information from Large Language Models: Syntactic Relations and Derivational KnowledgeTsedeniya Kinfe Temesgen, Marion Di Marco, Alexander FraserEMNLP 2025 · 被引用 2 次
