Correcting Large Language Model Behavior via Influence Function
Han Zhang, Zhuo Zhang, Yi Zhang, Yuanzhao Zhai, Hanyang Peng, Yu Lei, Yue Yu, Hui Wang, Bin Liang, Lin Gui, Ruifeng Xu
Abstract
Recent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate from contemporary human preferences and societal norms. Existing methodologies, either curation of new data for continual alignment or manual correction of outdated data for re-alignment, demand costly human resources. To address this, we propose a novel approach, LLM BehAvior Correction with INfluence FunCtion REcall and Post-Training (LANCET), which needs no human involvement. LANCET consists of two phases: (1) using a new method LinFAC to efficiently identify the training data that significantly impact undesirable model outputs, and (2) applying an novel Influence-driven Bregman Optimization (IBO) technique to adjust the model’s outputs based on these influence distributions. Our experiments show that LANCET effectively and efficiently corrects inappropriate behaviors of LLMs while preserving model utility. Further more, LANCET exhibits stronger generalization ability than all baselines under out-of-distribution harmful prompts, offering better interpretability and compatibility with real-world applications of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4436132d-c98f-400d-973d-ca84b18e6072Cited by top-tier papers5
- First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence EstimationDmytro Vitel, Anshuman ChhabraICLR 2026 · 8 citations
- Influence Functions for Edge Edits in Non-Convex Graph Neural NetworksJaeseung Heo, Kyeongheung Yun, Seokwon Yoon, MoonJeong Park et al.NeurIPS 2025 · 2 citations
- Enhancing Trustworthiness of Fine-Tuned LLMs via Regularized Subset SelectionKumar Shubham, Nishant Sharma, Karn Tiwari, Prathosh APICLR 2026
- The Realignment Problem: When Right becomes Wrong in LLMsAakash Sen Sharma, Debdeep Sanyal, Manodeep Ray, Vivek Srivastava et al.ICML 2026
- Beyond Binary Erasure: Soft-Weighted Unlearning for Fairness and RobustnessXinbao Qiao, Ningning Ding, Yushi Cheng, Meng ZhangAAAI 2026
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
Related papers
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 9 citations
- Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisKai Chen, Chunwei Wang, Kuo Yang, Jianhua Han et al.ICLR 2024 · 47 citations
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao et al.ICML 2023 · 287 citations
- How RLHF Amplifies SycophancyItai Shapira, Gerdus Benade, Ariel ProcacciaICML 2026 · 16 citations
- Mission Impossible: A Statistical Perspective on Jailbreaking LLMsJingtong Su, Julia Kempe, Karen UllrichNeurIPS 2024 · 38 citations
