Superbizarre Is Not Superb: Derivational Morphology Improves BERT's Interpretation of Complex Words
Valentin Hofmann, Janet B. Pierrehumbert, Hinrich Schütze
Abstract
How does the input segmentation of pretrained language models (PLMs) affect their interpretations of complex words? We present the first study investigating this question, taking BERT as the example PLM and focusing on its semantic representations of English derivatives. We show that PLMs can be interpreted as serial dual-route models, i.e., the meanings of complex words are either stored or else need to be computed from the subwords, which implies that maximally meaningful input tokens should allow for the best generalization on new words. This hypothesis is confirmed by a series of semantic probing tasks on which Del-BERT (Derivation leveraging BERT), a model with derivational input segmentation, substantially outperforms BERT with WordPiece segmentation. Our results suggest that the generalization capabilities of PLMs could be further improved if a morphologically-informed vocabulary of input tokens were used.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d655540-b7d6-40dc-8455-8738bb0b73f0Cited by top-tier papers16
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 301 citations
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai et al.NeurIPS 2024 · 23 citations
- Broken Tokens? Your Language Model can Secretly Handle Non-Canonical TokenizationsBrian Siyuan Zheng, Alisa Liu, Orevaoghene Ahia, Jonathan Hayase et al.NeurIPS 2025 · 19 citations
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelLeonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai et al.EMNLP 2023 · 10 citations
Builds on6
- CamemBERT: a Tasty French Language ModelLouis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont et al.ACL 2020 · 703 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection EncodingSamson Tan, Shafiq R. Joty, Lav R. Varshney, Min-Yen KanEMNLP 2020 · 26 citations
- Probing Pretrained Language Models for Lexical SemanticsIvan Vulic, Edoardo Maria Ponti, Robert Litschko, Goran Glavas et al.EMNLP 2020 · 26 citations
- Predicting the Growth of Morphological Families from Social and Linguistic FactorsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2020 · 13 citations
Related papers
- DagoBERT: Generating Derivational Morphology with a Pretrained Language ModelValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeEMNLP 2020 · 1 citation
- BERT-like Models for Slavic Morpheme SegmentationDmitry Morozov, Lizaveta Astapenka, Anna V. Glazkova, Timur Garipov et al.ACL 2025 · 1 citation
- Exploring morphology-aware tokenization: A case study on Spanish language modelingAlba Táboas García, Piotr Przybyla, Leo WannerEMNLP 2025 · 1 citation
- Improving Tokenisation by Alternative Treatment of SpacesEdward Gow-Smith, Harish Tayyar Madabushi, Carolina Scarton, Aline VillavicencioEMNLP 2022 · 6 citations
- BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model PerformanceTimo Schick, Hinrich SchützeACL 2020 · 1 citation
