Lune

ICLR2026Top-tier venue

Token Distillation: Attention-Aware Input Embeddings for New Tokens

Konstantin Dobler, Desmond Elliott, Gerard de Melo

2026Year
7Citations

Abstract

Current language models rely on static vocabularies determined at pretraining time, which can lead to decreased performance and increased computational cost for domains underrepresented in the original vocabulary. New tokens can be added to solve this problem, when coupled with a good initialization for their new embeddings. However, existing embedding initialization methods require expensive further training or pretraining of additional modules. In this paper, we propose Token Distillation and show that by distilling representations obtained using the original tokenization, we can quickly learn high-quality input embeddings for new tokens. Experimental results with a wide range of open-weight models show that Token Distillation outperforms even strong baselines. 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 43b449b2-d567-4677-b8dd-ff66d08bfca7

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines