Lune

KDD2026顶会

SCOPE: Streaming Covariance-Orthogonal Post-Hoc Editing for Continual LLM Safety Governance

Yizhe Yang, Xuanming Jiang, Jisheng Dang, Aoying Wang, Baoyi An, Hao Wu, Bimei Wang, Hong Peng, Guoshuai Zhao, Bin Hu, Zhongyu Yang

2026年份
1被引次数

摘要

Deployed large language models (LLMs) require continual post-hoc safety updates to counter evolving threats, yet current methods often entangle refusal behaviors with capability-critical representations, causing cumulative utility drift in sequential edits. We propose SCOPE, a streaming post-hoc editing framework for capability-invariant LLM safety governance. SCOPE derives layer-wise capability subspaces via online activation covariance and constrains safety updates to their orthogonal spaces, decoupling capability preservation from safety-specific edits. Within the derived null space, it extracts low-rank refusal directions from contrastive harmful activations to steer refusal behaviors without interfering with core capabilities. Its online covariance maintenance enables efficient, adaptive expansion of the protected capability subspace against open-ended threat evolution. Our experiments validate that SCOPE delivers significant safety robustness gains by substantially reducing attack success rates, while keeping general capabilities within natural run-to-run variance under repeated edits, supporting long-term safety maintenance for deployed LLMs.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get d998ce53-98ad-49e4-a5ec-e1c39bc24c72

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖