Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
Zhe Li, Wei Zhao, Yige Li, Jun Sun
摘要
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their deployment is frequently undermined by undesirable behaviors such as generating harmful content, factual inaccuracies, and societal biases. Diagnosing the root causes of these failures poses a critical challenge for AI safety. Existing attribution methods, particularly those based on parameter gradients, often fall short due to prohibitive noisy signals and computational complexity. In this work, we introduce a novel and efficient framework that diagnoses a range of undesirable LLM behaviors by analyzing representation and its gradients, which operates directly in the model's activation space to provide a semantically meaningful signal linking outputs to their training data. We systematically evaluate our method for tasks that include tracking harmful content, detecting backdoor poisoning, and identifying knowledge contamination. The results demonstrate that our approach not only excels at sample-level attribution but also enables fine-grained tokenlevel analysis, precisely identifying the specific samples and phrases that causally influence model behavior. This work provides a powerful diagnostic tool to understand, audit, and ultimately mitigate the risks associated with LLMs. The code is available at https://github.com/plumprc/RepT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora 等ICML 2024 · 被引用 460 次
相关 Paper
- Detecting and Filtering Unsafe Training Data via Data Attribution with Denoised RepresentationYijun Pan, Taiwei Shi, Jieyu Zhao, Jiaqi MaICML 2026 · 被引用 10 次
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMsGuangyu Shen, Siyuan Cheng, Xiangzhe Xu, Yuan Zhou 等ICML 2026
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang 等ICML 2026
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang 等ACL 2025
- LLMScan: Causal Scan for LLM Misbehavior DetectionMengdi Zhang, Kai Kiat Goh, Peixin Zhang, Jun Sun 等ICML 2025
