Token Distillation: Attention-Aware Input Embeddings for New Tokens
Konstantin Dobler, Desmond Elliott, Gerard de Melo
摘要
Current language models rely on static vocabularies determined at pretraining time, which can lead to decreased performance and increased computational cost for domains underrepresented in the original vocabulary. New tokens can be added to solve this problem, when coupled with a good initialization for their new embeddings. However, existing embedding initialization methods require expensive further training or pretraining of additional modules. In this paper, we propose Token Distillation and show that by distilling representations obtained using the original tokenization, we can quickly learn high-quality input embeddings for new tokens. Experimental results with a wide range of open-weight models show that Token Distillation outperforms even strong baselines. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 被引用 57 次
- Probing Pretrained Language Models for Lexical SemanticsIvan Vulic, Edoardo Maria Ponti, Robert Litschko, Goran Glavas 等EMNLP 2020 · 被引用 26 次
- Do All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai 等EMNLP 2023 · 被引用 24 次
相关 Paper
- TokAlign: Efficient Vocabulary Adaptation via Token AlignmentChong Li, Jiajun Zhang, Chengqing ZongACL 2025 · 被引用 7 次
- LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language ModelsHaoyu Wang, Ruirui Li, Haoming Jiang, Zhengyang Wang 等KDD 2023 · 被引用 5 次
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 被引用 48 次
- FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual ModelsKonstantin Dobler, Gerard de MeloEMNLP 2023 · 被引用 4 次
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio 等ACL 2023
