A Compact Model for Mathematics Problem Representations Distilled from BERT
Hao Ming, Xinguo Yu, Xiaotian Cheng, Zhenquan Shen, Xiaopan Lyu
Abstract
Large language models (LLMs) have made significant advancements in math problem solving, but their large size and high latency render them impractical for real-world applications in intelligent mathematics solvers. Recently, task-agnostic compact models have been developed to replace LLMs in general natural language processing tasks. However, these models often struggle to acquire sufficient math-related knowledge from LLMs, leading to unsatisfactory performance in solving math word problems (MWPs). To develop a specialized compact model for representing MWPs, we develop the knowledge distillation (KD) technique to extract mathematical semantics knowledge from the large pre-trained model BERT. Effective knowledge types and distillation strategies are explored through extensive experiments. Our KD algorithm employs multi-knowledge distillation to extract fundamental knowledge from hidden states in the middle to lower layers, while also incorporating knowledge of mathematical relations and symbol constraints from higher-layer outputs and math decoder outputs, by leveraging bottleneck networks. Pre-training tasks on MWP datasets, such as masked language modeling and part-of-speech tagging, are also utilized to enhance the generalization of the compact model for MWP understanding. Additionally, a simple parameter mixing strategy is employed to prevent catastrophic forgetting of acquired knowledge. Our findings indicate that our approach can reduce the size of a BERT model by 10% while retaining approximately 95% of its performance on MWP datasets, outperforming the mainstream BERT-based task-agnostic compact models. The efficacy of each component has been validated through ablation studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- HMS: A Hierarchical Solver with Dependency-Enhanced Understanding for Math Word ProblemXin Lin, Zhenya Huang, Hongke Zhao, Enhong Chen et al.AAAI 2021 · 70 citations
- Semantically-Aligned Universal Tree-Structured Solver for Math Word ProblemsJinghui Qin, Lihui Lin, Xiaodan Liang, Rumin Zhang et al.EMNLP 2020 · 62 citations
Related papers
- XtremeDistil: Multi-stage Distillation for Massive Multilingual ModelsSubhabrata Mukherjee, Ahmed Hassan AwadallahACL 2020 · 4 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Distilling Linguistic Context for Language Model CompressionGeondo Park, Gyeongman Kim, Eunho YangEMNLP 2021 · 24 citations
