TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
Felipe Pinto Coelho Nuti, Tim Franzmeyer, João F. Henriques
摘要
Past work has studied the effects of fine-tuning on large language models' (LLMs) overall performance on certain tasks. However, a way to quantitatively analyze its effect on individual outputs is still lacking. In this work, we propose a new method for measuring the contribution that fine-tuning makes to individual LLM responses using the model's intermediate hidden states, and assuming access to the original pre-trained model. We introduce and theoretically analyze an exact decomposition of any fine-tuned LLM into a pretraining component and a fine-tuning component. Empirically, we find that one can steer model behavior and performance by up-or down-scaling the fine-tuning component during the forward pass. Motivated by this finding and our theoretical analysis, we define the Tuning Contribution (TuCo) in terms of the ratio of the fine-tuning component and the pre-training component. We find that three prominent adversarial attacks on LLMs circumvent safety measures in a way that reduces the Tuning Contribution, and that TuCo is consistently lower on prompts where the attacks succeed compared to ones where they do not. This suggests that attenuating the effect of fine-tuning on model outputs plays a role in the success of these attacks. In short, TuCo enables the quantitative study of how fine-tuning influences model behavior and safety, and vice-versa. 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
相关 Paper
- Token-level Data Selection for Safe LLM Fine-tuningYanping Li, Zhening Liu, Zijian Li, Zehong Lin 等ICLR 2026 · 被引用 4 次
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 被引用 7 次
- Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse DatasetsNing Lu, Shengcai Liu, Jiahao Wu, Weiyu Chen 等ICML 2025
- RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data SelectionYixin Yang, Qingxiu Dong, Linli Yao, Fangwei Zhu 等AAAI 2026
- Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuningGuoli Wang, Haonan Shi, Tu Ouyang, An WangKDD 2026 · 被引用 5 次
