FIRST: Teach A Reliable Large Language Model Through Efficient Trustworthy Distillation
KaShun Shum, Minrui Xu, Jianshu Zhang, Zixin Chen, Shizhe Diao, Hanze Dong, Jipeng Zhang, Muhammad Omer Raza
摘要
Large language models (LLMs) have become increasingly prevalent in our daily lives, leading to an expectation for LLMs to be trustworthy -both accurate and well-calibrated (the prediction confidence should align with its ground truth correctness likelihood). Nowadays, fine-tuning has become the most popular method for adapting a model to practical usage by significantly increasing accuracy on downstream tasks. Despite the great accuracy it achieves, we found fine-tuning is still far away from satisfactory trustworthiness due to "tuning-induced mis-calibration". In this paper, we delve deeply into why and how miscalibration exists in fine-tuned models, and how distillation can alleviate the issue. Then we further propose a brand new method named EFfIcient TRustworthy DiSTillation (FIRST), which utilizes a small portion of teacher's knowledge to obtain a reliable language model in a cost-efficient way. Specifically, we identify the "concentrated knowledge" phenomenon during distillation, which can significantly reduce the computational burden. Then we apply a "trustworthy maximization" process to optimize the utilization of this small portion of concentrated knowledge before transferring it to the student. Experimental results demonstrate the effectiveness of our method, where better accuracy (+2.3%) and less mis-calibration (-10%) are achieved on average across both indomain and out-of-domain scenarios, indicating better trustworthiness. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity OptimizationYichen Yan, Ming Zhong, Qi Zhu, Xiaoling Gu 等NeurIPS 2025 · 被引用 8 次
- Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMsAnshumann, Mohd Abbas Zaidi, Akhil Kedia, Jinwoo Ahn 等ACL 2025
它引用的顶会 Paper7
- On-Policy Distillation of Language Models: Learning from Self-Generated MistakesRishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk 等ICLR 2024 · 被引用 311 次
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 143 次
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 被引用 102 次
- Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-StepLiunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren 等ACL 2023 · 被引用 34 次
- Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusTianhang Zhang, Lin Qiu, Qipeng Guo, Cheng Deng 等EMNLP 2023 · 被引用 18 次
相关 Paper
- Logits-Based FinetuningJingyao Li, Senqiao Yang, Sitong Wu, Han Shi 等EMNLP 2025
- Calibrating LLMs with Information-Theoretic Evidential Deep LearningYawei Li, David Rügamer, Bernd Bischl, Mina RezaeiICLR 2025
- DUET: Distilled LLM Unlearning from an Efficiently Contextualized TeacherYisheng Zhong, Zhengbang Yang, Zhuangdi ZhuICLR 2026 · 被引用 4 次
- PELA: Learning Parameter-Efficient Models with Low-Rank ApproximationYangyang Guo, Guangzhi Wang, Mohan S. KankanhalliCVPR 2024
- Towards Objective Fine-tuning: How LLMs' Prior Knowledge Causes Potential Poor Calibration?Ziming Wang, Zeyu Shi, Haoyi Zhou, Shiqi Gao 等ACL 2025 · 被引用 6 次
