Residual Stream Analysis of Overfitting And Structural Disruptions
Quan Liu, Han Zhou, Wenquan Wu, Hua Wu, Sen Su
摘要
Ensuring that large language models (LLMs) remain both helpful and harmless poses a significant challenge: fine-tuning on repetitive safety datasets-where unsafe prompts are paired with standard refusal templates-often leads to false refusals, in which benign queries are declined. We first quantify this effect, showing that safety data exhibits substantially lower token entropy (H 1 ≈ 9.18) and 2gram diversity (≈ 0.048) compared to general instruction data (H 1 ≈ 12.05, 2-gram≈0.205). To uncover the root cause, we introduce FlowLens, a stable PCA-based tool for residual-stream geometry analysis, and reveal that higher proportions of safety examples concentrate variance along a few components, reducing representational smoothness and driving false refusals (false refusal rate rises from 63% to 84% as safety data increases from 0% to 40%). Guided by these insights, we propose Variance Concentration Loss (VCL), an auxiliary regularizer that penalizes excessive variance concentration in mid-layer residuals. Empirical results demonstrate that VCL reduces false refusals by over 35 percentage points while maintaining or improving performance on general benchmarks such as MMLU and GSM8K.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- WildChat: 1M ChatGPT Interaction Logs in the WildWenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 等ICLR 2024 · 被引用 504 次
相关 Paper
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation SteeringUtsav Maskey, Sumit Yadav, Mark Dras, Usman NaseemACL 2026
- Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector AblationXinpeng Wang, Chengzhi Hu, Paul Röttger, Barbara PlankICLR 2025 · 被引用 1 次
- Safety Anchor: Defending Harmful Fine-tuning via Geometric BottlenecksGuoxin Lu, Letian Sha, Qing Wang, Peijie Sun 等ICML 2026
- ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMsZijun Chen, Wenbo Hu, Ya Li, Lei Miao 等ICLR 2026
- Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive DecodingYupeng Qi, Ziyu Lyu, Lixin Cui, Lu Bai 等ACL 2026
