Enhancing Trustworthiness of Fine-Tuned LLMs via Regularized Subset Selection
Kumar Shubham, Nishant Sharma, Karn Tiwari, Prathosh AP
摘要
Supervised fine-tuning (SFT) improves large language model (LLM) perplexity, but can also degrade trustworthiness-leading to the generation of untruthful, biased, or unsafe content during user interactions. These issues are often traced back to specific phrases or patterns in the training data. However, correcting them usually requires expensive retraining or new data collection. In this work, we propose a two-stage, compute-efficient repair of the post-SFT models that enhances trustworthiness while preserving the downstream performance. In the first stage, we identify the training samples responsible for failures on trustworthiness metrics like truthfulness, stereotypical bias, and machine ethics-and select a small, diverse subset of these examples using a determinantal point process (DPP)-based regularization. In the second stage, we repair the model under the framework of proximal Bregman response function (PBRF) using a gradient ascent update, which enhances trustworthiness while preserving downstream task performance (perplexity). We evaluate our method on multiple LLMs of varying sizes and demonstrate up to 21% improvement in trustworthiness metrics with minimal impact (≤ 1%) on perplexity. Our method provides a computationally efficient approach to enhance post-SFT models and offers a practical alternative to hours of retraining required for model repair. Our code is available at https://github.com/kyrs/tracing-llm-trust . * Equal Contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint TuningWei Chen, Zhen Huang, Liang Xie, Binbin Lin 等ICML 2024 · 被引用 55 次
- More RLHF, More Trust? On The Impact of Preference Alignment On TrustworthinessAaron Jiaxun Li, Satyapriya Krishna, Himabindu LakkarajuICLR 2025
- Anchored Supervised Fine-TuningHe Zhu, Junyou Su, Peng Lai, Ren Ma 等ICLR 2026 · 被引用 13 次
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment QualityYuto Harada, Yusuke Yamauchi, Yusuke Oda, Yohei Oseki 等EMNLP 2025
- Reasoning Quality Emerges Early: Data Curation for Reasoning ModelsHongyi Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato 等ICML 2026
