Structural Knowledge Distillation: Tractably Distilling Information for Structured Predictor
Xinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia, Nguyen Bach, Tao Wang, Zhongqiang Huang, Fei Huang, Kewei Tu
摘要
Knowledge distillation is a critical technique to transfer knowledge between models, typically from a large model (the teacher) to a more fine-grained one (the student). The objective function of knowledge distillation is typically the cross-entropy between the teacher and the student's output distributions. However, for structured prediction problems, the output space is exponential in size; therefore, the cross-entropy objective becomes intractable to compute and optimize directly. In this paper, we derive a factorized form of the knowledge distillation objective for structured prediction, which is tractable for many typical choices of the teacher and student models. In particular, we show the tractability and empirical effectiveness of structural knowledge distillation between sequence labeling and dependency parsing models under four different scenarios: 1) the teacher and student share the same factorization form of the output structure scoring function; 2) the student factorization produces more fine-grained substructures than the teacher factorization; 3) the teacher factorization produces more fine-grained substructures than the student factorization; 4) the factorization forms from the teacher and the student are incompatible. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Neural Span-Based Continual Named Entity Recognition ModelYunan Zhang, Qingcai ChenAAAI 2023 · 被引用 16 次
- Measuring and Reducing Model Update Regression in Structured Prediction for NLPDeng Cai, Elman Mansimov, Yi-An Lai, Yixuan Su 等NeurIPS 2022 · 被引用 14 次
- Language Modelling via Learning to RankArvid Frydenlund, Gagandeep Singh, Frank RudziczAAAI 2022 · 被引用 9 次
- LogicMP: A Neuro-symbolic Approach for Encoding First-order Logic ConstraintsWeidi Xu, Jingwei Wang, Lele Xie, Jianshan He 等ICLR 2024 · 被引用 6 次
它引用的顶会 Paper4
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Efficient Second-Order TreeCRF for Neural Dependency ParsingYu Zhang, Zhenghua Li, Min ZhangACL 2020 · 被引用 90 次
- Structure-Level Knowledge Distillation For Multilingual Sequence LabelingXinyu Wang, Yong Jiang, Nguyen Bach, Tao Wang 等ACL 2020 · 被引用 31 次
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 被引用 4 次
相关 Paper
- f-Divergence Minimization for Sequence-Level Knowledge DistillationYuqiao Wen, Zichao Li, Wenyu Du, Lili MouACL 2023 · 被引用 14 次
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 被引用 7 次
- Progressively Knowledge Distillation via Re-parameterizing Diffusion Reverse ProcessXufeng Yao, Fanbin Lu, Yuechen Zhang, Xinyun Zhang 等AAAI 2024 · 被引用 7 次
- Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency ParsingZhuoran Li, Chunming Hu, Junfan Chen, Zhijun Chen 等AAAI 2025 · 被引用 1 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
