Lune

ACL2023Top-tier venue

f-Divergence Minimization for Sequence-Level Knowledge Distillation

Yuqiao Wen, Zichao Li, Wenyu Du, Lili Mou

2023Year
14Citations
46Top-tier citations

Abstract

Knowledge distillation (KD) is the process of transferring knowledge from a large model to a small one. It has gained increasing attention in the natural language processing community, driven by the demands of compressing evergrowing language models. In this work, we propose an f -DISTILL framework, which formulates sequence-level knowledge distillation as minimizing a generalized f -divergence function. We propose four distilling variants under our framework and show that existing SeqKD and ENGINE approaches are approximations of our f -DISTILL methods. We further derive step-wise decomposition for our f -DISTILL, reducing intractable sequence-level divergence to word-level losses that can be computed in a tractable manner. Experiments across four datasets show that our methods outperform existing KD approaches, and that our symmetric distilling losses can better force the student to learn from the teacher distribution. 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext db8f0f55-d315-43ee-bdad-27cbcc273735

Cited by top-tier papers46

Ask how each one uses it

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines