SRA: Span Representation Alignment for Large Language Model Distillation
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Trung Le
摘要
Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ different tokenizers. While existing approaches mainly focus on token-level alignment strategies, which are often brittle and sensitive to discrepancies between tokenizers, we argue that the method of aggregating tokens into more robust representations before distillation is of equal importance. In this paper, we introduce SRA (Span Representation Alignment for Large Language Model Distillation), a novel framework that reframes CTKD through the physical lens of Multi-Particle Dynamical Systems. SRA shifts the fundamental unit of alignment from tokens to robust, tokenizer-agnostic spans. We model each span as a cluster of particles and represent its state by its Center of Mass (CoM) - an attention-weighted average that captures rich semantic information. We leverage the concept of span centers of mass with attention-derived weighting to prioritize the most salient spans. In addition, we employ a geometric regularizer to preserve the structural integrity of the representation space and introduce aligned span logit distillation to enhance knowledge transfer across models. In challenging cross-architecture distillation experiments, SRA consistently and significantly outperforms state-of-the-art CTKD baselines, validating our physically-grounded approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi 等EMNLP 2022 · 被引用 238 次
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan 等ICLR 2024 · 被引用 113 次
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
- DistiLLM: Towards Streamlined Distillation for Large Language ModelsJongwoo Ko, Sungnyun Kim, Tianyi Chen, Se-Young YunICML 2024 · 被引用 86 次
相关 Paper
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie 等AAAI 2025 · 被引用 31 次
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep 等EMNLP 2025
- Entropy-aware Span-Constrained Optimal Transport for Robust Cross-Tokenizer Knowledge DistillationZhi-Ping Liu, Simiao Li, Wei Li, Hanting Chen 等ICML 2026
- Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer DistillationHuazheng Wang, Yongcheng Jing, Haifeng Sun, Jingyu Wang 等AAAI 2026
- MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language ModelsHoang Tran Vuong, Tue Le, Quyen Tran, Linh Ngo Van 等AAAI 2026
