Token-Scaled Logit Distillation for Ternary Weight Generative Language Models
Minsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong, Du-Seong Chang, Wonyong Sung, Jungwook Choi
Abstract
Generative Language Models (GLMs) have shown impressive performance in tasks such as text generation, understanding, and reasoning. However, the large model size poses challenges for practical deployment. To solve this problem, Quantization-Aware Training (QAT) has become increasingly popular. However, current QAT methods for generative models have resulted in a noticeable loss of accuracy. To counteract this issue, we propose a novel knowledge distillation method specifically designed for GLMs. Our method, called token-scaled logit distillation, prevents overfitting and provides superior learning from the teacher model and ground truth. This research marks the first evaluation of ternary weight quantization-aware training of large-scale GLMs with less than 1.0 degradation in perplexity and achieves enhanced accuracy in tasks like common-sense QA and arithmetic reasoning as well as natural language understanding. 2 * Corresponding Author 2 Our code is available at https://github.com/aiha-lab/TSLD 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Accurate Retraining-free Pruning for Pretrained Encoder-based Language ModelsSeungcheol Park, Hojun Choi, U KangICLR 2024 · 14 citations
- Improving Conversational Abilities of Quantized Large Language Models via Direct Preference AlignmentJanghwan Lee, Seongmin Park, Sukjin Hong, Minsoo Kim et al.ACL 2024 · 2 citations
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling ModelsJiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo et al.ICLR 2026
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- TernaryBERT: Distillation-aware Ultra-low Bit BERTWei Zhang, Lu Hou, Yichun Yin, Lifeng Shang et al.EMNLP 2020 · 147 citations
- Compression of Generative Pre-trained Language Models via QuantizationChaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang et al.ACL 2022 · 119 citations
Related papers
- BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-DistillationDayou Du, Yijia Zhang, Shijie Cao, Jiaqi Guo et al.ACL 2024 · 16 citations
- LFQ: Logit-aware Final-block Quantization for Boosting the Generation Quality of Low-Bit Quantized LLMsJung Hyun Lee, June Yong Yang, Jungwook Choi, Eunho YangICML 2026
- Understanding and Improving Knowledge Distillation for Quantization Aware Training of Large Transformer EncodersMinsoo Kim, Sihwa Lee, Sukjin Hong, Du-Seong Chang et al.EMNLP 2022 · 7 citations
- LLM-Oriented Token-Adaptive Knowledge DistillationXurong Xie, Zhucun Xue, Jiafu Wu, Jian Li et al.AAAI 2026
- Gated Relational Alignment via Confidence-based Distillation for Efficient VLMsYanlong Chen, Amir Habibian, Luca Benini, Yawei LiICML 2026 · 5 citations
