Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language Models
Xiao Cui, Mo Zhu, Yulei Qin, Liang Xie, Wengang Zhou, Houqiang Li
Abstract
Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers (i.e., vocabularies) between teacher and student models, limiting their versatility in handling LLMs of different architecture families. In this paper, we introduce the Multi-Level Optimal Transport (MultiLevelOT), a novel approach that advances the optimal transport for universal cross-tokenizer knowledge distillation. Our method aligns the logit distributions of the teacher and the student at both token and sequence levels using diverse cost matrices, eliminating the need for dimensional or token-by-token correspondence. At the token level, MultiLevelOT integrates both global and local information by jointly optimizing all tokens within a sequence to enhance robustness. At the sequence level, we efficiently capture complex distribution structures of logits via the Sinkhorn distance, which approximates the Wasserstein distance for divergence measures. Extensive experiments on tasks such as extractive QA, generative QA, and summarization demonstrate that the MultiLevelOT outperforms state-of-the-art cross-tokenizer KD methods under various settings. Our approach is robust to different student and teacher models across model families, architectures, and parameter sizes. Codes and models are available at https: //github.com/2018cx/Multi-Level-OT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 716b6f14-d2c6-4010-b2dd-11cc74027898Cited by top-tier papers18
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model DistillationBuu Phan, Ashish Khisti, Karen UllrichICLR 2026 · 5 citations
- Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset DistillationXiao Cui, Yulei Qin, Wengang Zhou, Hongsheng Li et al.NeurIPS 2025 · 5 citations
- Knowledge Distillation for Large Language Models through Residual LearningThinh On, Hengzhi Pei, Leonard Lausen, George KarypisICLR 2026 · 5 citations
- SRA: Span Representation Alignment for Large Language Model DistillationQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen et al.ACL 2026 · 1 citation
- Explainable Token-level Noise Filtering for LLM Fine-tuning DatasetsYuchen Yang, Wenze Lin, Enhao Huang, Zhixuan Chu et al.ICLR 2026 · 1 citation
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk et al.ICLR 2024 · 311 citations
- Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie et al.ACL 2022 · 131 citations
Related papers
- MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language ModelsHoang Tran Vuong, Tue Le, Quyen Tran, Linh Ngo Van et al.AAAI 2026
- Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge DistillationThong Thanh Nguyen, Anh Tuan LuuAAAI 2022 · 46 citations
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
- Entropy-aware Span-Constrained Optimal Transport for Robust Cross-Tokenizer Knowledge DistillationZhi-Ping Liu, Simiao Li, Wei Li, Hanting Chen et al.ICML 2026
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
