Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer Distillation
Huazheng Wang, Yongcheng Jing, Haifeng Sun, Jingyu Wang, Jianxin Liao, Leszek Rutkowski, Dacheng Tao
Abstract
Cross-tokenizer knowledge distillation, where the teacher and student employ different tokenizers, is becoming increasingly prevalent, yet it poses underexplored challenges: existing methods fail to capture the rich knowledge encoded in teacher logits, as evidenced by the neglect of semantic information, inaccurate and biased logit alignment, and discarding distributional structure—ultimately leading to unfavorable distillation. To address these issues, we propose SeDi, a semantics and distribution-aware knowledge transfer framework tailored for cross-tokenizer distillation. To preserve factual knowledge, SeDi employs bipartite graph-based alignment at the tokenization level and a sliding window re-encoding strategy at the vocabulary level, enabling unbiased transfer of the teacher’s next-token predictions into the student’s vocabulary space. To further retain distributional information, we align the student’s entropy with that of the teacher by incorporating the student’s own logits during training, which helps to mitigate the exposure bias problem. Experiments on ten datasets across three task domains and five different teacher-student model pairs with varying vocabulary sizes demonstrate that SeDi delivers substantial improvements, with gains of up to 19.8%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8575f230-16a8-4b6a-9d40-2fed7626437cBuilds on16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
Related papers
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
- Entropy-aware Span-Constrained Optimal Transport for Robust Cross-Tokenizer Knowledge DistillationZhi-Ping Liu, Simiao Li, Wei Li, Hanting Chen et al.ICML 2026
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
- SRA: Span Representation Alignment for Large Language Model DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen et al.ACL 2026 · 1 citation
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 5 citations
